Close the learning loop
You run a decision, record that an expert overrode it, and ask the platform for a policy change that would have agreed with the expert. The platform replays the proposal against your golden cases and passes it only with zero regressions. You admit it explicitly as a new version, check the new behavior, and roll back.
About this run
The decision here is a renewal discount policy: account managers may discount up to 20%, and anything deeper goes to a VP. The examples were run on 2026-09-28 against a local instance of the reasoning service built from its current source, which includes the fixes listed under Known issues; you run them against your reasoning API endpoint. How the loop fits the platform is described in How learning works.
What you'll build#
- Admit a playbook as version 1 of your pack.
- Decide two renewals with it.
- Record what happened: one expert override, one success.
- Ask for a policy proposal and read the golden regression gate.
- Admit the proposal as version 2, check it, and roll back to version 1.
Prerequisites#
- Hosted API access: your reasoning API endpoint in
HELIXOR_API_URLand your API key inHELIXOR_API_KEY. - Python with
httpxinstalled. - Two people, or at least two identities: admission requires a reviewer who is not the principal your API key belongs to.
What the loop does in this release#
| Stage | Endpoint | What it does |
|---|---|---|
| Decide | POST /api/v1/reasoning/playbooks/turn | Evaluates the active version of a pack against a typed state. Returns a verdict, the move, the rule IDs and a run ID. |
| Observe | POST /api/v1/reasoning/decisions/{run_id}/outcomes | Records an outcome and classifies it as a data defect, a policy defect or expected performance. |
| Propose and gate | POST /api/v1/reasoning/playbooks/evolve | Scores candidate changes against the recorded policy defects and your golden cases. Returns a proposal only if the gate passes. Changes nothing. |
| Admit | POST /api/v1/reasoning/playbooks/admit | Validates, runs a preflight simulation, and stores the pack as a new version that becomes active. |
| Roll back | POST /api/v1/reasoning/playbooks/{pack_id}/versions/{version}/activate | Makes an earlier version active again, with a reason. |
Two facts shape everything below. Proposals are threshold changes: each numeric comparison in a hard rule is moved by −5%, −2%, +2% and +5%. And the gate is strict: a proposal must resolve at least one recorded gap and cause zero golden regressions.
Steps#
- Write the pack and a small client
A hosted playbook states its rules as conditions over typed state. This pack has two hard rules and three disposition bands:
{ "pack_id": "renewals.discount_guard.v1", "goal_id": "renewals.discount_guard.v1", "kind": "business", "description": "Renewal offers: rep discount ceiling and sponsor escalation for large accounts.", "floor": 0.80, "budget": {"max_latency_ms": 30.0}, "actions": [ "offer_standard_renewal", "apply_discount_tier", "schedule_exec_sponsor", "escalate_to_vp" ], "hard_rules": [ { "id": "RULE-MAX-DISCOUNT-20", "description": "Rep discount ceiling is 20%; above it the VP decides", "condition": "state.discount_pct > 20.0 and state.user_role != 'vp_sales'", "action": "escalate_to_vp", "verdict": "refuse" }, { "id": "RULE-LARGE-ACCOUNT", "description": "Accounts above 250k a year get an executive sponsor", "condition": "state.annual_revenue > 250000", "action": "schedule_exec_sponsor", "verdict": "queue" } ], "disposition_bands": { "act": {"min_p": 0.80, "verdict": "answer", "rule_id": "DISP-RENEWAL"}, "close": {"max_p": 0.39, "verdict": "refuse", "rule_id": "DISP-CHURN"}, "queue": {"min_p": 0.40, "max_p": 0.79, "verdict": "queue", "rule_id": "DISP-REVIEW"} } }The client wraps the API and fails on any error status:
import os import httpx API = httpx.Client( base_url=os.environ["HELIXOR_API_URL"], # your reasoning API endpoint headers={"X-API-Key": os.environ["HELIXOR_API_KEY"]}, timeout=10.0, ) PACK_ID = "renewals.discount_guard.v1" def call(method, path, **kwargs): resp = API.request(method, path, **kwargs) if resp.status_code >= 400: raise RuntimeError(f"{method} {path} -> {resp.status_code}: {resp.text}") return resp.json() def decide(state, message="Renewal offer"): d = call("POST", "/api/v1/reasoning/playbooks/turn", json={"pack_id": PACK_ID, "message": message, "state": state}) return d - Admit the pack as version 1
Admission puts the pack under version control. The reviewer is recorded with the version and must be a different identity from the principal your API key belongs to.
import json from loop_client import call pack = json.load(open("renewal_pack.json")) r = call("POST", "/api/v1/reasoning/playbooks/admit", json={ "pack": pack, "reviewed_by": "pricing-lead@example.com", "review_note": "Initial policy", }) for key in ("admitted", "pack_id", "version", "reviewed_by", "admitted_by", "content_sha256", "receipt_hash"): print(f"{key:15}: {r[key]}")admitted : True pack_id : renewals.discount_guard.v1 version : 1 reviewed_by : pricing-lead@example.com admitted_by : anonymous content_sha256 : a6a6544bd2a6445ad8eb8471416bd329cd80db57c34e0b1f8b12d056dcde03c8 receipt_hash : 8cf36c41c17c8c5b06174fb25328fbb07a9ebe942e245cf27ec6d6f72221c0aa
admitted_byis the principal behind your API key; the local instance ran without sign-in, so it readsanonymous.content_sha256is the digest of the pack's canonical JSON (keys sorted, compact), so the order you write keys in does not change it.receipt_hashis the receipt of the preflight simulation that admission ran; it differs on every admission. Admitting with yourself as reviewer is refused:{"detail":{"failure_code":"playbook_self_review","detail":"The reviewer must be a different person from the admitting principal.","diagnostics":[]}} HTTP 403 - Decide two renewals
import json from loop_client import decide cases = { "renewal-4410": {"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000}, "renewal-4411": {"discount_pct": 12.0, "user_role": "account_manager", "annual_revenue": 60000}, } runs = {} for name, state in cases.items(): d = decide(state) runs[name] = {"run_id": d["workflow_run"]["run_id"], "state": state, "move": d["chosen_move"]} print(f"{name}: verdict={d['verdict']:7} move={d['chosen_move']:22} rules={d['rule_ids']}") print(f" run_id={d['workflow_run']['run_id']} receipt={d['proof']['receipt_hash']}") print(f" pack v{d['pack_version']} ({d['pack_source']}) digest={d['pack_digest'][:16]}...") json.dump(runs, open("runs.json", "w"), indent=2)renewal-4410: verdict=refuse move=escalate_to_vp rules=['RULE-MAX-DISCOUNT-20'] run_id=wf-1790601908457 receipt=98fde62c0291c0d59a3ad667 pack v1 (admitted) digest=a6a6544bd2a6445a... renewal-4411: verdict=answer move=offer_standard_renewal rules=['DISP-RENEWAL'] run_id=wf-1790601908468 receipt=227f1d392c3e11bf0c518610 pack v1 (admitted) digest=a6a6544bd2a6445a...The 21% discount breaches the 20% ceiling and escalates to a VP. The 12% discount passes. Keep
workflow_run.run_idwith your record: outcomes are recorded against it. Keeppack_versionandpack_digesttoo: they name the policy that decided.pack_sourceisadmittedfor a version from the registry,builtinfor a pack that ships with the service andlocalfor a pack you supply that is not in the registry, andpack_versionisnullfor the last two. - Record what happened
The VP approved the 21% discount, so the pack's escalation was overridden. Report that as an
overridewith the action the expert chose, and put the state the decision saw inobserved_value: the proposal step replays it. Report the other renewal as asuccess.import json from loop_client import PACK_ID, call runs = json.load(open("runs.json")) # The VP approved the 21% discount the pack escalated, and the customer renewed. r = runs["renewal-4410"] o1 = call("POST", f"/api/v1/reasoning/decisions/{r['run_id']}/outcomes", json={ "pack_id": PACK_ID, "outcome_label": "override", "decision_move": r["move"], "expert_override_action": "apply_discount_tier", "observed_value": r["state"], # the state the decision saw "human_notes": "VP approved 21%; one point over the ceiling is routine for multi-year renewals", }) # The standard renewal was accepted as offered. r = runs["renewal-4411"] o2 = call("POST", f"/api/v1/reasoning/decisions/{r['run_id']}/outcomes", json={ "pack_id": PACK_ID, "outcome_label": "success", "decision_move": r["move"], "observed_value": r["state"], }) for o in (o1, o2): print(f"{o['run_id']}: {o['outcome_label']:8} attribution={o['attribution']:21} " f"gap_example_id={o['gap_example_id']}") gaps = call("GET", "/api/v1/reasoning/decisions/outcomes", params={"pack_id": PACK_ID, "attribution": "policy_defect"}) print(f"\npolicy defects on record for {PACK_ID}: {len(gaps)}")wf-1790601908457: override attribution=policy_defect gap_example_id=gap-wf-1790601908457 wf-1790601908468: success attribution=expected_performance gap_example_id=None policy defects on record for renewals.discount_guard.v1: 1
The service classified each outcome. An
override,failure,default,disputeorcancelledoutcome is a data defect when thedata_qualityyou send with it is low or stale, and otherwise a policy defect. Anything else is expected performance. Only policy defects and overrides feed proposals; a data defect is fixed at the source, not in the policy. You can also setattributionyourself. - Ask for a proposal and read the gate
Golden cases are decisions the policy must keep getting right whatever changes. Pass them with the request, or declare them in the pack as
golden_vectors. With neither, the service refuses to propose anything:422 {"detail":{"failure_code":"GOLDEN_SET_REQUIRED","message":"Pack 'renewals.discount_guard.v1' declares no golden_vectors and the request supplied none; evolution needs a golden set to gate regressions."}}With
pack_id, the service uses the active version and every policy defect recorded for that pack.import json from loop_client import PACK_ID, call # Cases the policy must keep getting right, whatever changes. golden = [ {"name": "small_discount", "expected_verdict": "answer", "state": {"discount_pct": 10.0, "user_role": "account_manager", "annual_revenue": 60000}}, {"name": "deep_discount", "expected_verdict": "refuse", "state": {"discount_pct": 30.0, "user_role": "account_manager", "annual_revenue": 60000}}, {"name": "vp_deep_discount", "expected_verdict": "answer", "state": {"discount_pct": 30.0, "user_role": "vp_sales", "annual_revenue": 60000}}, {"name": "large_account", "expected_verdict": "queue", "state": {"discount_pct": 5.0, "user_role": "account_manager", "annual_revenue": 400000}}, ] report = call("POST", "/api/v1/reasoning/playbooks/evolve", json={"pack_id": PACK_ID, "golden_vectors": golden}) print(f"gap examples : {report['gap_examples_evaluated']}") print(f"golden cases : {report['golden_cases_checked']}") print(f"mutations : {report['mutations_evaluated']}") print(f"gate passed : {report['golden_regression_gate_passed']}\n") print(f"{'candidate':58} resolved regressions admissible") for m in report["pareto_candidates"]: print(f"{m['description']:58} {m['gap_cases_resolved']:8} {m['golden_regressions']:11} {m['is_admissible']}") champion = report["admitted_champion"] print("\nchampion :", champion and champion["mutation_id"]) print("new condition:", champion and champion["new_condition"]) for line in report["diagnostics"]: print("diagnostic :", line) if champion: json.dump(champion["candidate_pack"], open("proposal.json", "w"), indent=2)gap examples : 1 golden cases : 4 mutations : 8 gate passed : True candidate resolved regressions admissible Adjust threshold discount_pct from 20.0 to 21.0 1 0 True Adjust threshold discount_pct from 20.0 to 19.0 0 0 False Adjust threshold discount_pct from 20.0 to 19.6 0 0 False Adjust threshold discount_pct from 20.0 to 20.4 0 0 False Adjust threshold annual_revenue from 250000.0 to 237500.0 0 0 False champion : mut-renewals.discount_guard.v1-4 new condition: state.discount_pct > 21.0 and state.user_role != 'vp_sales' diagnostic : NO_EXCEPTION_TEMPLATES: the pack declares no exception_templates; only threshold mutations were proposed. diagnostic : EVOLUTION_SUCCESS: Champion 'mut-renewals.discount_guard.v1-4' resolved 1/1 gap cases with ZERO golden suite regressions. Proposal only: admit it through the reviewed admission gate to make it live.
Eight candidates were scored, four per hard rule; the report lists the top five. Only raising the ceiling to 21% resolves the override (21 is not greater than 21.0), and it breaks none of the four golden cases, so the gate passes and it is the proposal. Raising it to 20.4% keeps the golden cases but does not resolve the override, so it is not admissible. Candidates are ranked by admissibility, then gaps resolved, then fewer regressions. The
NO_EXCEPTION_TEMPLATESdiagnostic says the pack declares no exception clauses the service may propose, so only thresholds were tried; exception templates, like golden cases, come from the pack or the request.Now add one more golden case, a 20.5% discount that must still escalate, and run the same request:
{"name": "just_over_ceiling", "expected_verdict": "refuse", "state": {"discount_pct": 20.5, "user_role": "account_manager", "annual_revenue": 60000}},gap examples : 1 golden cases : 5 mutations : 8 gate passed : False candidate resolved regressions admissible Adjust threshold discount_pct from 20.0 to 21.0 1 1 False Adjust threshold discount_pct from 20.0 to 19.0 0 0 False Adjust threshold discount_pct from 20.0 to 19.6 0 0 False Adjust threshold discount_pct from 20.0 to 20.4 0 0 False Adjust threshold annual_revenue from 250000.0 to 237500.0 0 0 False champion : None new condition: None diagnostic : NO_EXCEPTION_TEMPLATES: the pack declares no exception_templates; only threshold mutations were proposed. diagnostic : EVOLUTION_NO_CHAMPION: None of the candidate mutations passed the zero-regression golden gate while resolving gap cases.
The 21% candidate still resolves the override, but it now lets a 20.5% discount through without a VP, which is a regression. No candidate passes, so there is no proposal. This is the gate doing its job: your golden cases, not the override alone, decide what may change.
- Admit the proposal as version 2
A passing proposal is still only a proposal. Nothing changes until someone admits it. Run step 4 again without the extra golden case so
proposal.jsonholds the passing candidate, then:import json from loop_client import PACK_ID, call, decide proposal = json.load(open("proposal.json")) r = call("POST", "/api/v1/reasoning/playbooks/admit", json={ "pack": proposal, "reviewed_by": "pricing-lead@example.com", "review_note": "Raise rep ceiling to 21% after VP overrides", }) print(f"admitted version {r['version']} of {r['pack_id']} (preflight receipt {r['receipt_hash'][:16]}...)") d = decide({"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000}) print(f"21% now: verdict={d['verdict']} move={d['chosen_move']} rules={d['rule_ids']}") v = call("GET", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions") print(f"\nversions={v['versions']} active={v['active_version']}") for e in v["audit"]: print(f" {e['event']:9} v{e['version']} by={e['actor']} reviewed_by={e['reviewed_by']}")admitted version 2 of renewals.discount_guard.v1 (preflight receipt faff28decac86ca4...) 21% now: verdict=answer move=offer_standard_renewal rules=['DISP-RENEWAL'] versions=[1, 2] active=2 admitted v1 by=anonymous reviewed_by=pricing-lead@example.com admitted v2 by=anonymous reviewed_by=pricing-lead@example.com
The same 21% renewal no longer escalates. Note what the change does and does not do: it removes the escalation, so the decision falls through to the
DISP-RENEWALband and returns the pack's first action,offer_standard_renewal. It does not learn to choose the VP'sapply_discount_tier; to select that action, add a rule for it and admit that version. - Roll back
from loop_client import PACK_ID, call, decide r = call("POST", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions/1/activate", params={"reason": "Finance asked to restore the 20% ceiling"}) print(r) d = decide({"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000}) print(f"21% now: verdict={d['verdict']} move={d['chosen_move']} rules={d['rule_ids']}") v = call("GET", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions") print(f"versions={v['versions']} active={v['active_version']}") e = v["audit"][-1] print(f" {e['event']:9} v{e['version']} previous=v{e['previous_active_version']} reason={e['reason']!r}"){'pack_id': 'renewals.discount_guard.v1', 'active_version': 1, 'reviewed_by': 'pricing-lead@example.com'} 21% now: verdict=refuse move=escalate_to_vp rules=['RULE-MAX-DISCOUNT-20'] versions=[1, 2] active=1 activated v1 previous=v2 reason='Finance asked to restore the 20% ceiling'Activation never creates or edits a version; version 2 is still stored, and the audit log records who activated what, when, why, and what was active before.
Known issues#
Fixed on main: the turn route takes the tenant from your credentials
Service builds before 2026-09-28 read tenant_id for POST /api/v1/reasoning/playbooks/turn from the request body (default local), so a caller could name any tenant. On main the tenant and user come from your credentials: omit tenant_id, or send your own tenant; a different tenant is refused with 403 and failure_code tenant_mismatch. Conversation memory, recorded outcomes and /playbooks/evolve are scoped to your tenant, and evolve requires credentials. On an older build, call the turn route only from your own backend and never pass a tenant from client input.
Fixed in the service's source on 2026-09-28
Earlier builds had these defects; an endpoint that has not been updated still has them:
- A stored version could become unreadable when
disposition_bandskeys were not in alphabetical order (500registry_content_hash_mismatchafter a restart or on rollback). Versions are now hashed over the same canonical JSON they are stored in, so key order does not matter. On an older build, list the keys alphabetically. - Omitting
golden_vectorsused two built-in cases written for a lending pack. Golden cases now come from the request or the pack, and without them the call is refused withGOLDEN_SET_REQUIRED. On an older build, always send your own. - Exception clauses came from one built-in template for lending. They now come from
exception_templatesin the pack or the request. - The turn response did not name the version that decided. It now carries
pack_version,pack_digestandpack_source. On an older build, readactive_versionfromGET /api/v1/reasoning/playbooks/{pack_id}/versions.
Still open: the service does not check that a run_id belongs to a decision it made. Use the one from the turn response.
How it works#
For each recorded policy defect, the service rebuilds the decision from observed_value and asks, for each candidate, whether any hard rule still fires. A gap counts as resolved when no hard rule fires and the expert's action is one of the pack's actions. Each candidate is also run through the preflight simulator against your golden cases; a case whose verdict or rule ID differs from what you expected is a regression. A condition that cannot be evaluated, for example because a state field is missing, counts as a refusal. The candidate with zero regressions and the most resolved gaps is proposed.
Admission runs the same validation and preflight simulation as any new pack, then writes the pack as a new immutable version with its content hash, the reviewer, the admitting principal and the preflight receipt, and makes it active. Versions, outcomes and the audit log are stored by the service and were still there after the local instance was restarted.
Limits#
- Proposals change thresholds of existing rules. The loop does not add rules, remove rules or change which action a rule selects; you do that in a new version.
- The registry is per service, not per tenant: admitting a version of a pack ID changes it for every caller of that pack ID on that service. Use pack IDs you own.
- The proposal is only as good as your golden cases. Cover the boundaries you care about, as the 20.5% case shows.
- Store
pack_versionandpack_digestfrom the turn response with each decision; the active version can change between two calls. - The embedded runtime does not run hosted playbooks.