Adaptive strategies
Use observations to choose which permitted strategy to try next. Define the objective, check each result, and learn which strategies work for which problem contexts while preserving the playbook's hard constraints.
Availability
Contextual strategy selection and conditional lowering are experimental. Automatic integration with public playbooks and evaluate() is planned. This guide describes how to design and evaluate that integration; it does not introduce callable APIs or supported configuration fields.
What you can use now#
| Capability | How you use it | Boundary |
|---|---|---|
| Embedded decision packs | Evaluate declared rules with the embedded runtime. | Evaluation does not choose strategies from outcome memory. |
| Standalone outcome memory | Record observations with HelixorBeliefLedger; your application connects it to decisions. | Weighted belief is not a calibrated probability of success. Check the tutorial's release notes before using persistence helpers. |
| Hosted policy learning, preview | Record outcomes, propose threshold changes, check golden cases, and explicitly admit a new version in the learning loop. | It does not automatically synthesize and execute adaptive strategies. |
| Adaptive strategy loop | Planned Rank admitted alternatives by context, execute within a budget, verify, and revise. | No public strategy-loop API or pack-schema extension is documented yet. |
Choose a problem with observable outcomes#
Start with structured state and several permitted ways to reach a measurable objective. Recovery, diagnostic sequencing, workflow repair and selection among supported solver strategies are candidates. A strategy can be useful even when it is only the best next experiment: a diagnostic may rule out a cause without completing the task.
Expect value when different contexts favor different strategies, unsuccessful attempts provide information, and another attempt fits the budget. Compare against a fixed strategy and an ordinary feedback loop. Learning can add cost without benefit when one strategy already dominates, feedback is unreliable, or the problem changes faster than evidence accumulates.
Keep language understanding separate. A proposed interpretation of a math problem is an assumption until validated. A solver can prove a result under that interpretation without proving that it represents the original question. Strategy learning cannot repair an unsupported representation by increasing its confidence score.
Example: repair first or diagnose first#
This is a conceptual example, not an executable playbook. Suppose a job failed and the objective is a completed job whose output passes an independent integrity check. The playbook permits a reversible retry, a dependency diagnostic and escalation. It forbids bypassing access checks or discarding input.
| Observed context | Candidate strategy | What to verify and learn |
|---|---|---|
| Transient interruption; dependencies healthy | Retry once, validate output, then escalate if still failing. | A passing output check establishes success for this run. Record the context and attempt cost. |
| Repeated failure; dependency state unknown | Diagnose first, then select a permitted repair using the observation. | A useful diagnosis is progress. Only a repaired job with valid output is final success. |
| Missing or contradictory diagnostic evidence | Request an allowed additional observation or escalate. | Record uncertainty. Do not relabel an unverified repair as successful. |
A ledger may eventually favor retry-first in one context and diagnose-first in another. It must remain possible to reconsider a low-ranked strategy when observations change. A low estimated success rate is a reason to defer an attempt; it is not proof that the strategy is impossible.
Define the loop before learning#
- Declare the objective and verifier
State what counts as completion, what evidence the verifier accepts, and which assumptions remain unresolved. Separate the executor's report from the evidence used to validate it.
- Admit legal alternatives
Validate typed state and strategy preconditions. Apply hard constraints before ranking. A learned preference cannot authorize a prohibited action or remove a hard rule.
- Rank by context and evidence
Use attributes available before the decision: problem type, structural motif, observed state and operating conditions. Keep strategy, pack, executor and verifier versions with the evidence. Define minimum evidence and behavior for unknown or stale contexts.
- Execute within a budget
Bound attempts, elapsed time and cost. Recheck preconditions before each action. Concurrent attempts need isolated state or explicit coordination so they cannot apply conflicting repairs; their correlated results are not independent evidence.
- Verify and record
Record the attempted strategy, observations, cost, progress and terminal outcome against the decision. An unavailable verifier leaves the outcome unverified. Deduplicate repeated outcome deliveries before updating estimates.
- Revise or stop
Use new observations to rerank permitted alternatives. Stop on verified completion, explicit failure, exhausted budget or an escalation condition. A new strategy outside the admitted set is a proposal requiring validation and admission.
Keep authoritative decision and outcome records as the source of evidence. Treat strategy estimates as derived views that can be rebuilt from those records, rather than introducing a second source of decision truth.
What the probability means#
Keep three claims separate: the interpretation is appropriate, the chosen strategy will succeed under the stated conditions, and the returned result passed a verifier. Each needs its own evidence. Do not multiply their scores as if they were independent or label a model-conditional result as an unconditional proof.
For success forecasting, count complete attempts against a declared objective. Preserve failures, timeouts and unverified outcomes as distinct observations; do not silently drop difficult cases. Track partial progress separately so several useful steps do not become several successful jobs. A unit-weight success/failure estimate has a different meaning from a belief score that deliberately penalizes failures more heavily.
Before exposing a probability to users, test calibration on held-out cases with a frozen ledger: among cases assigned similar probabilities, how often did the objective actually succeed? Report sample counts, uncertainty, coverage and error among accepted cases, plus cost and abstentions. Test each important context and changed operating conditions. Sparse evidence and distribution shift can make a confident estimate unreliable.
A threshold is not a guarantee
A predicted success threshold is a selection rule. It does not establish an upper bound on actual failures. Validate the accepted subset separately, including uncertainty and the conditions under which that validation applies.
Compose strategies from reusable fragments#
In the planned integration, a fragment describes a useful subprocedure: its entry conditions, effects, consumed resources, cost and source evidence. A successful local procedure can become a candidate building block for a larger problem. Expand and revalidate its original operations before execution, including obligations at the boundaries between fragments. Reuse does not transfer a previous result's certificate to a new problem.
Compare composition with an existing planner as well as flat random and genetic search. When a declared state model supports lookahead, ordinary planning may solve the problem more reliably and cheaply. Genetic search can propose new combinations; it should not be chosen merely because a problem is difficult. Count every primitive action inside a fragment when comparing search budgets.
Estimate the probability of the whole strategy from complete outcomes in a matching context. Two fragments can fail together because they share a dependency, so multiplying their individual success rates can misrepresent the combined strategy. Keep success forecasts separate from independent checks that a certified result answers the original question.
Scope evidence to relevant versions and operating conditions. An explicit version change can invalidate an estimate; an unobserved change can leave an old estimate dangerously confident. Specify how fresh outcomes trigger investigation, withdrawal of estimates and retraining. Do not promise that the probability threshold itself detects such a change.
What a risk control may change#
A planned risk control can select among admitted operating policies: spend more on diagnostics, require more evidence before acting, allow a bounded exploratory attempt, or escalate earlier. Keep effort budget and acceptance threshold explicit; spending less effort does not imply a known error rate.
Hard constraints, authorization, typed admission and proof requirements remain fixed. If a policy permits returning an unverified candidate, present it with its assumptions and evidence status. A control cannot turn “probably correct” into “verified.”
How this changes playbook authoring#
The compiled pack schema and hosted playbooks are different interfaces. Do not add fields from this design to either format. When adaptive integration is exposed, its versioned contract will need to define the following behavior:
| Authoring concern | Required decision |
|---|---|
| Objective and outcomes | Define completion, independent verification, partial progress and unresolved outcomes. |
| Alternatives and probes | Declare permitted strategies, preconditions, observations, costs and stopping rules. |
| Evidence context | Define attributes known at selection time, version scope, evidence floors and stale-evidence handling. |
| Selection semantics | Align ranking, ties, abstention, reported forecasts and conditions for reconsidering a committed plan. Changing a selector alone is insufficient. |
| Policy changes | Separate choosing among admitted alternatives from editing rules or adding strategies. Changes to policy still require review and admission. |
| Replay | Capture typed input, pack version, selector and executor versions, ledger snapshot, external observations and any random seed. |
Updating evidence can legitimately change the next strategy without changing the policy. Replay must freeze that evidence as well as the input. This planned behavior does not change the embedded runtime's existing evaluation contract.
Prove that it helps#
Use separate training, calibration and test cases. Compare the fixed baseline, feedback without a learned ledger, and contextual strategy selection on the same cases with equal execution budgets. Keep test outcomes out of the ledger during evaluation. Measure verified completions, failures, cost, attempts and unresolved cases; retain paired wins and losses so an average cannot hide a harmed context.
Add cases for cold starts, misleading observations, duplicate outcomes, partial repairs, exhausted budgets, changed conditions and exact replay. Evaluate exploration on its own: observations exist only for strategies you actually attempted unless a simulator provides counterfactual outcomes, and simulator evidence does not establish production performance. See Testing an adaptive integration.
Start with the runnable outcome-memory tutorial to learn the existing ledger interface. Use the hosted learning loop for reviewed policy proposals. Automatic strategy execution remains planned.