helixordevelopers

Close the learning loop

You run a decision, record that an expert overrode it, and ask the platform for a policy change that would have agreed with the expert. The platform replays the proposal against your golden cases and passes it only with zero regressions. You admit it explicitly as a new version, check the new behavior, and roll back.

PreviewHosted API40 minutesAdvancedPython 3.9+

About this run

The decision here is a renewal discount policy: account managers may discount up to 20%, and anything deeper goes to a VP. The examples were run on 2026-09-28 against a local instance of the reasoning service built from its current source, which includes the fixes listed under Known issues; you run them against your reasoning API endpoint. How the loop fits the platform is described in How learning works.

What you'll build#

  1. Admit a playbook as version 1 of your pack.
  2. Decide two renewals with it.
  3. Record what happened: one expert override, one success.
  4. Ask for a policy proposal and read the golden regression gate.
  5. Admit the proposal as version 2, check it, and roll back to version 1.

Prerequisites#

  • Hosted API access: your reasoning API endpoint in HELIXOR_API_URL and your API key in HELIXOR_API_KEY.
  • Python with httpx installed.
  • Two people, or at least two identities: admission requires a reviewer who is not the principal your API key belongs to.

What the loop does in this release#

StageEndpointWhat it does
DecidePOST /api/v1/reasoning/playbooks/turnEvaluates the active version of a pack against a typed state. Returns a verdict, the move, the rule IDs and a run ID.
ObservePOST /api/v1/reasoning/decisions/{run_id}/outcomesRecords an outcome and classifies it as a data defect, a policy defect or expected performance.
Propose and gatePOST /api/v1/reasoning/playbooks/evolveScores candidate changes against the recorded policy defects and your golden cases. Returns a proposal only if the gate passes. Changes nothing.
AdmitPOST /api/v1/reasoning/playbooks/admitValidates, runs a preflight simulation, and stores the pack as a new version that becomes active.
Roll backPOST /api/v1/reasoning/playbooks/{pack_id}/versions/{version}/activateMakes an earlier version active again, with a reason.

Two facts shape everything below. Proposals are threshold changes: each numeric comparison in a hard rule is moved by −5%, −2%, +2% and +5%. And the gate is strict: a proposal must resolve at least one recorded gap and cause zero golden regressions.

Steps#

  1. Write the pack and a small client

    A hosted playbook states its rules as conditions over typed state. This pack has two hard rules and three disposition bands:

    {
      "pack_id": "renewals.discount_guard.v1",
      "goal_id": "renewals.discount_guard.v1",
      "kind": "business",
      "description": "Renewal offers: rep discount ceiling and sponsor escalation for large accounts.",
      "floor": 0.80,
      "budget": {"max_latency_ms": 30.0},
      "actions": [
        "offer_standard_renewal",
        "apply_discount_tier",
        "schedule_exec_sponsor",
        "escalate_to_vp"
      ],
      "hard_rules": [
        {
          "id": "RULE-MAX-DISCOUNT-20",
          "description": "Rep discount ceiling is 20%; above it the VP decides",
          "condition": "state.discount_pct > 20.0 and state.user_role != 'vp_sales'",
          "action": "escalate_to_vp",
          "verdict": "refuse"
        },
        {
          "id": "RULE-LARGE-ACCOUNT",
          "description": "Accounts above 250k a year get an executive sponsor",
          "condition": "state.annual_revenue > 250000",
          "action": "schedule_exec_sponsor",
          "verdict": "queue"
        }
      ],
      "disposition_bands": {
        "act":   {"min_p": 0.80, "verdict": "answer", "rule_id": "DISP-RENEWAL"},
        "close": {"max_p": 0.39, "verdict": "refuse", "rule_id": "DISP-CHURN"},
        "queue": {"min_p": 0.40, "max_p": 0.79, "verdict": "queue", "rule_id": "DISP-REVIEW"}
      }
    }
    

    The client wraps the API and fails on any error status:

    import os
    
    import httpx
    
    API = httpx.Client(
        base_url=os.environ["HELIXOR_API_URL"],          # your reasoning API endpoint
        headers={"X-API-Key": os.environ["HELIXOR_API_KEY"]},
        timeout=10.0,
    )
    PACK_ID = "renewals.discount_guard.v1"
    
    
    def call(method, path, **kwargs):
        resp = API.request(method, path, **kwargs)
        if resp.status_code >= 400:
            raise RuntimeError(f"{method} {path} -> {resp.status_code}: {resp.text}")
        return resp.json()
    
    
    def decide(state, message="Renewal offer"):
        d = call("POST", "/api/v1/reasoning/playbooks/turn",
                 json={"pack_id": PACK_ID, "message": message, "state": state})
        return d
    
  2. Admit the pack as version 1

    Admission puts the pack under version control. The reviewer is recorded with the version and must be a different identity from the principal your API key belongs to.

    import json
    
    from loop_client import call
    
    pack = json.load(open("renewal_pack.json"))
    r = call("POST", "/api/v1/reasoning/playbooks/admit", json={
        "pack": pack,
        "reviewed_by": "pricing-lead@example.com",
        "review_note": "Initial policy",
    })
    for key in ("admitted", "pack_id", "version", "reviewed_by", "admitted_by", "content_sha256", "receipt_hash"):
        print(f"{key:15}: {r[key]}")
    
    admitted       : True
    pack_id        : renewals.discount_guard.v1
    version        : 1
    reviewed_by    : pricing-lead@example.com
    admitted_by    : anonymous
    content_sha256 : a6a6544bd2a6445ad8eb8471416bd329cd80db57c34e0b1f8b12d056dcde03c8
    receipt_hash   : 8cf36c41c17c8c5b06174fb25328fbb07a9ebe942e245cf27ec6d6f72221c0aa
    

    admitted_by is the principal behind your API key; the local instance ran without sign-in, so it reads anonymous. content_sha256 is the digest of the pack's canonical JSON (keys sorted, compact), so the order you write keys in does not change it. receipt_hash is the receipt of the preflight simulation that admission ran; it differs on every admission. Admitting with yourself as reviewer is refused:

    {"detail":{"failure_code":"playbook_self_review","detail":"The reviewer must be a different person from the admitting principal.","diagnostics":[]}}
    HTTP 403
    
  3. Decide two renewals
    import json
    
    from loop_client import decide
    
    cases = {
        "renewal-4410": {"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000},
        "renewal-4411": {"discount_pct": 12.0, "user_role": "account_manager", "annual_revenue": 60000},
    }
    runs = {}
    for name, state in cases.items():
        d = decide(state)
        runs[name] = {"run_id": d["workflow_run"]["run_id"], "state": state, "move": d["chosen_move"]}
        print(f"{name}: verdict={d['verdict']:7} move={d['chosen_move']:22} rules={d['rule_ids']}")
        print(f"              run_id={d['workflow_run']['run_id']}  receipt={d['proof']['receipt_hash']}")
        print(f"              pack v{d['pack_version']} ({d['pack_source']})  digest={d['pack_digest'][:16]}...")
    
    json.dump(runs, open("runs.json", "w"), indent=2)
    
    renewal-4410: verdict=refuse  move=escalate_to_vp         rules=['RULE-MAX-DISCOUNT-20']
                  run_id=wf-1790601908457  receipt=98fde62c0291c0d59a3ad667
                  pack v1 (admitted)  digest=a6a6544bd2a6445a...
    renewal-4411: verdict=answer  move=offer_standard_renewal rules=['DISP-RENEWAL']
                  run_id=wf-1790601908468  receipt=227f1d392c3e11bf0c518610
                  pack v1 (admitted)  digest=a6a6544bd2a6445a...
    

    The 21% discount breaches the 20% ceiling and escalates to a VP. The 12% discount passes. Keep workflow_run.run_id with your record: outcomes are recorded against it. Keep pack_version and pack_digest too: they name the policy that decided. pack_source is admitted for a version from the registry, builtin for a pack that ships with the service and local for a pack you supply that is not in the registry, and pack_version is null for the last two.

  4. Record what happened

    The VP approved the 21% discount, so the pack's escalation was overridden. Report that as an override with the action the expert chose, and put the state the decision saw in observed_value: the proposal step replays it. Report the other renewal as a success.

    import json
    
    from loop_client import PACK_ID, call
    
    runs = json.load(open("runs.json"))
    
    # The VP approved the 21% discount the pack escalated, and the customer renewed.
    r = runs["renewal-4410"]
    o1 = call("POST", f"/api/v1/reasoning/decisions/{r['run_id']}/outcomes", json={
        "pack_id": PACK_ID,
        "outcome_label": "override",
        "decision_move": r["move"],
        "expert_override_action": "apply_discount_tier",
        "observed_value": r["state"],        # the state the decision saw
        "human_notes": "VP approved 21%; one point over the ceiling is routine for multi-year renewals",
    })
    
    # The standard renewal was accepted as offered.
    r = runs["renewal-4411"]
    o2 = call("POST", f"/api/v1/reasoning/decisions/{r['run_id']}/outcomes", json={
        "pack_id": PACK_ID,
        "outcome_label": "success",
        "decision_move": r["move"],
        "observed_value": r["state"],
    })
    
    for o in (o1, o2):
        print(f"{o['run_id']}: {o['outcome_label']:8} attribution={o['attribution']:21} "
              f"gap_example_id={o['gap_example_id']}")
    
    gaps = call("GET", "/api/v1/reasoning/decisions/outcomes",
                params={"pack_id": PACK_ID, "attribution": "policy_defect"})
    print(f"\npolicy defects on record for {PACK_ID}: {len(gaps)}")
    
    wf-1790601908457: override attribution=policy_defect         gap_example_id=gap-wf-1790601908457
    wf-1790601908468: success  attribution=expected_performance  gap_example_id=None
    
    policy defects on record for renewals.discount_guard.v1: 1
    

    The service classified each outcome. An override, failure, default, dispute or cancelled outcome is a data defect when the data_quality you send with it is low or stale, and otherwise a policy defect. Anything else is expected performance. Only policy defects and overrides feed proposals; a data defect is fixed at the source, not in the policy. You can also set attribution yourself.

  5. Ask for a proposal and read the gate

    Golden cases are decisions the policy must keep getting right whatever changes. Pass them with the request, or declare them in the pack as golden_vectors. With neither, the service refuses to propose anything:

    422 {"detail":{"failure_code":"GOLDEN_SET_REQUIRED","message":"Pack 'renewals.discount_guard.v1' declares no golden_vectors and the request supplied none; evolution needs a golden set to gate regressions."}}
    

    With pack_id, the service uses the active version and every policy defect recorded for that pack.

    import json
    
    from loop_client import PACK_ID, call
    
    # Cases the policy must keep getting right, whatever changes.
    golden = [
        {"name": "small_discount", "expected_verdict": "answer",
         "state": {"discount_pct": 10.0, "user_role": "account_manager", "annual_revenue": 60000}},
        {"name": "deep_discount", "expected_verdict": "refuse",
         "state": {"discount_pct": 30.0, "user_role": "account_manager", "annual_revenue": 60000}},
        {"name": "vp_deep_discount", "expected_verdict": "answer",
         "state": {"discount_pct": 30.0, "user_role": "vp_sales", "annual_revenue": 60000}},
        {"name": "large_account", "expected_verdict": "queue",
         "state": {"discount_pct": 5.0, "user_role": "account_manager", "annual_revenue": 400000}},
    ]
    
    report = call("POST", "/api/v1/reasoning/playbooks/evolve",
                  json={"pack_id": PACK_ID, "golden_vectors": golden})
    
    print(f"gap examples : {report['gap_examples_evaluated']}")
    print(f"golden cases : {report['golden_cases_checked']}")
    print(f"mutations    : {report['mutations_evaluated']}")
    print(f"gate passed  : {report['golden_regression_gate_passed']}\n")
    print(f"{'candidate':58} resolved  regressions  admissible")
    for m in report["pareto_candidates"]:
        print(f"{m['description']:58} {m['gap_cases_resolved']:8}  {m['golden_regressions']:11}  {m['is_admissible']}")
    
    champion = report["admitted_champion"]
    print("\nchampion     :", champion and champion["mutation_id"])
    print("new condition:", champion and champion["new_condition"])
    for line in report["diagnostics"]:
        print("diagnostic   :", line)
    
    if champion:
        json.dump(champion["candidate_pack"], open("proposal.json", "w"), indent=2)
    
    gap examples : 1
    golden cases : 4
    mutations    : 8
    gate passed  : True
    
    candidate                                                  resolved  regressions  admissible
    Adjust threshold discount_pct from 20.0 to 21.0                   1            0  True
    Adjust threshold discount_pct from 20.0 to 19.0                   0            0  False
    Adjust threshold discount_pct from 20.0 to 19.6                   0            0  False
    Adjust threshold discount_pct from 20.0 to 20.4                   0            0  False
    Adjust threshold annual_revenue from 250000.0 to 237500.0         0            0  False
    
    champion     : mut-renewals.discount_guard.v1-4
    new condition: state.discount_pct > 21.0 and state.user_role != 'vp_sales'
    diagnostic   : NO_EXCEPTION_TEMPLATES: the pack declares no exception_templates; only threshold mutations were proposed.
    diagnostic   : EVOLUTION_SUCCESS: Champion 'mut-renewals.discount_guard.v1-4' resolved 1/1 gap cases with ZERO golden suite regressions. Proposal only: admit it through the reviewed admission gate to make it live.
    

    Eight candidates were scored, four per hard rule; the report lists the top five. Only raising the ceiling to 21% resolves the override (21 is not greater than 21.0), and it breaks none of the four golden cases, so the gate passes and it is the proposal. Raising it to 20.4% keeps the golden cases but does not resolve the override, so it is not admissible. Candidates are ranked by admissibility, then gaps resolved, then fewer regressions. The NO_EXCEPTION_TEMPLATES diagnostic says the pack declares no exception clauses the service may propose, so only thresholds were tried; exception templates, like golden cases, come from the pack or the request.

    Now add one more golden case, a 20.5% discount that must still escalate, and run the same request:

        {"name": "just_over_ceiling", "expected_verdict": "refuse",
         "state": {"discount_pct": 20.5, "user_role": "account_manager", "annual_revenue": 60000}},
    
    gap examples : 1
    golden cases : 5
    mutations    : 8
    gate passed  : False
    
    candidate                                                  resolved  regressions  admissible
    Adjust threshold discount_pct from 20.0 to 21.0                   1            1  False
    Adjust threshold discount_pct from 20.0 to 19.0                   0            0  False
    Adjust threshold discount_pct from 20.0 to 19.6                   0            0  False
    Adjust threshold discount_pct from 20.0 to 20.4                   0            0  False
    Adjust threshold annual_revenue from 250000.0 to 237500.0         0            0  False
    
    champion     : None
    new condition: None
    diagnostic   : NO_EXCEPTION_TEMPLATES: the pack declares no exception_templates; only threshold mutations were proposed.
    diagnostic   : EVOLUTION_NO_CHAMPION: None of the candidate mutations passed the zero-regression golden gate while resolving gap cases.
    

    The 21% candidate still resolves the override, but it now lets a 20.5% discount through without a VP, which is a regression. No candidate passes, so there is no proposal. This is the gate doing its job: your golden cases, not the override alone, decide what may change.

  6. Admit the proposal as version 2

    A passing proposal is still only a proposal. Nothing changes until someone admits it. Run step 4 again without the extra golden case so proposal.json holds the passing candidate, then:

    import json
    
    from loop_client import PACK_ID, call, decide
    
    proposal = json.load(open("proposal.json"))
    r = call("POST", "/api/v1/reasoning/playbooks/admit", json={
        "pack": proposal,
        "reviewed_by": "pricing-lead@example.com",
        "review_note": "Raise rep ceiling to 21% after VP overrides",
    })
    print(f"admitted version {r['version']} of {r['pack_id']} (preflight receipt {r['receipt_hash'][:16]}...)")
    
    d = decide({"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000})
    print(f"21% now: verdict={d['verdict']} move={d['chosen_move']} rules={d['rule_ids']}")
    
    v = call("GET", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions")
    print(f"\nversions={v['versions']} active={v['active_version']}")
    for e in v["audit"]:
        print(f"  {e['event']:9} v{e['version']}  by={e['actor']}  reviewed_by={e['reviewed_by']}")
    
    admitted version 2 of renewals.discount_guard.v1 (preflight receipt faff28decac86ca4...)
    21% now: verdict=answer move=offer_standard_renewal rules=['DISP-RENEWAL']
    
    versions=[1, 2] active=2
      admitted  v1  by=anonymous  reviewed_by=pricing-lead@example.com
      admitted  v2  by=anonymous  reviewed_by=pricing-lead@example.com
    

    The same 21% renewal no longer escalates. Note what the change does and does not do: it removes the escalation, so the decision falls through to the DISP-RENEWAL band and returns the pack's first action, offer_standard_renewal. It does not learn to choose the VP's apply_discount_tier; to select that action, add a rule for it and admit that version.

  7. Roll back
    from loop_client import PACK_ID, call, decide
    
    r = call("POST", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions/1/activate",
             params={"reason": "Finance asked to restore the 20% ceiling"})
    print(r)
    
    d = decide({"discount_pct": 21.0, "user_role": "account_manager", "annual_revenue": 90000})
    print(f"21% now: verdict={d['verdict']} move={d['chosen_move']} rules={d['rule_ids']}")
    
    v = call("GET", f"/api/v1/reasoning/playbooks/{PACK_ID}/versions")
    print(f"versions={v['versions']} active={v['active_version']}")
    e = v["audit"][-1]
    print(f"  {e['event']:9} v{e['version']}  previous=v{e['previous_active_version']}  reason={e['reason']!r}")
    
    {'pack_id': 'renewals.discount_guard.v1', 'active_version': 1, 'reviewed_by': 'pricing-lead@example.com'}
    21% now: verdict=refuse move=escalate_to_vp rules=['RULE-MAX-DISCOUNT-20']
    versions=[1, 2] active=1
      activated v1  previous=v2  reason='Finance asked to restore the 20% ceiling'
    

    Activation never creates or edits a version; version 2 is still stored, and the audit log records who activated what, when, why, and what was active before.

Known issues#

Fixed on main: the turn route takes the tenant from your credentials

Service builds before 2026-09-28 read tenant_id for POST /api/v1/reasoning/playbooks/turn from the request body (default local), so a caller could name any tenant. On main the tenant and user come from your credentials: omit tenant_id, or send your own tenant; a different tenant is refused with 403 and failure_code tenant_mismatch. Conversation memory, recorded outcomes and /playbooks/evolve are scoped to your tenant, and evolve requires credentials. On an older build, call the turn route only from your own backend and never pass a tenant from client input.

Fixed in the service's source on 2026-09-28

Earlier builds had these defects; an endpoint that has not been updated still has them:

  • A stored version could become unreadable when disposition_bands keys were not in alphabetical order (500 registry_content_hash_mismatch after a restart or on rollback). Versions are now hashed over the same canonical JSON they are stored in, so key order does not matter. On an older build, list the keys alphabetically.
  • Omitting golden_vectors used two built-in cases written for a lending pack. Golden cases now come from the request or the pack, and without them the call is refused with GOLDEN_SET_REQUIRED. On an older build, always send your own.
  • Exception clauses came from one built-in template for lending. They now come from exception_templates in the pack or the request.
  • The turn response did not name the version that decided. It now carries pack_version, pack_digest and pack_source. On an older build, read active_version from GET /api/v1/reasoning/playbooks/{pack_id}/versions.

Still open: the service does not check that a run_id belongs to a decision it made. Use the one from the turn response.

How it works#

For each recorded policy defect, the service rebuilds the decision from observed_value and asks, for each candidate, whether any hard rule still fires. A gap counts as resolved when no hard rule fires and the expert's action is one of the pack's actions. Each candidate is also run through the preflight simulator against your golden cases; a case whose verdict or rule ID differs from what you expected is a regression. A condition that cannot be evaluated, for example because a state field is missing, counts as a refusal. The candidate with zero regressions and the most resolved gaps is proposed.

Admission runs the same validation and preflight simulation as any new pack, then writes the pack as a new immutable version with its content hash, the reviewer, the admitting principal and the preflight receipt, and makes it active. Versions, outcomes and the audit log are stored by the service and were still there after the local instance was restarted.

Limits#

  • Proposals change thresholds of existing rules. The loop does not add rules, remove rules or change which action a rule selects; you do that in a new version.
  • The registry is per service, not per tenant: admitting a version of a pack ID changes it for every caller of that pack ID on that service. Use pack IDs you own.
  • The proposal is only as good as your golden cases. Cover the boundaries you care about, as the 20.5% case shows.
  • Store pack_version and pack_digest from the turn response with each decision; the active version can change between two calls.
  • The embedded runtime does not run hosted playbooks.

Next steps#