Track outcomes and confidence
A decision is only as good as what happens after it. You record each outcome against the rule that made the decision, per context such as a customer segment, and read back a confidence that your application uses to route the next decision: automatic, sampled, reviewed, or switched off.
Status
Outcome memory ships in the runtime as HelixorBeliefLedger. It is standalone: evaluate() does not read or write it, so you connect the two in your code, as this tutorial shows. Its interface is not yet covered by the API reference and may change. Feeding outcomes back into a new policy version is the job of the hosted learning loop; see Close the learning loop.
What you'll build#
- A short script that shows how belief moves with each confirmation, contradiction and partial outcome.
- A routing function for a refund rule,
refunds.auto_approve_under_50, that decides per customer segment whether its decisions run automatically, run with sampling, go to review, or stop. - A snapshot file that restores the ledger after a restart.
Prerequisites#
- The runtime installed. See Installation. No license, account or network access is needed.
- A decision whose outcome you learn later: a refund that was or was not charged back, an approval that was or was not overridden.
The model#
The ledger keeps, for each claim and each context, two tallies: confirmations c and contradictions x. Belief is the mean of a Beta posterior with a uniform prior, where each contradiction counts pain multiplier times (3 by default):
belief = (1 + c) / ((1 + c) + (1 + 3·x))
| Call | Effect |
|---|---|
record(claim, "CONFIRM", strength=s) | Adds s to c. |
record(claim, "CONTRADICT", strength=s) | Adds s to x. |
record(claim, "PARTIAL", strength=s) | With s ≤ 1, adds s to c and 1 − s to x. With s > 1, adds s/2 to each. |
belief(claim, context=None) | The formula above; 0.5 for a claim or context never observed. Without context, it pools every observation of the claim. |
is_bedrock(claim, context=None) | True when c + x ≥ 20 and belief ≥ 0.90. |
is_dead(claim, context=None) | True when c + x ≥ 5 and belief ≤ 0.15. |
snapshot() Next release | The whole ledger as a JSON-serializable helixor.belief_ledger.v1 document. |
HelixorBeliefLedger.restore(snapshot) Next release | A new ledger from a snapshot. Raises BeliefSnapshotError for a wrong schema_version, missing or unknown keys, or counts that are negative or not finite. |
evidence_mass, posterior, variance, claims() Next release | The evidence c + x, the Beta parameters (1 + c, 1 + 3·x), the posterior variance, and every claim with observations. |
The evidence counted by is_bedrock and is_dead is the plain sum c + x, without the multiplier. The thresholds are fixed in HelixorBeliefLedger; the multiplier is set with HelixorBeliefLedger(pain_multiplier=...). outcome is case-insensitive. An unknown outcome, a negative strength or a multiplier that is not positive raises ValueError.
Steps#
- See how belief moves
from helixor_runtime import HelixorBeliefLedger ledger = HelixorBeliefLedger() claim = "refunds.auto_approve_under_50" print("unobserved :", round(ledger.belief(claim), 3)) for _ in range(4): ledger.record(claim, "CONFIRM", context="segment:returning") print("4 confirms :", round(ledger.belief(claim, context="segment:returning"), 3)) ledger.record(claim, "CONTRADICT", context="segment:returning") print("4 confirms, 1 contra :", round(ledger.belief(claim, context="segment:returning"), 3)) ledger.record(claim, "PARTIAL", strength=0.5, context="segment:returning") print("+ PARTIAL 0.5 :", round(ledger.belief(claim, context="segment:returning"), 3)) print("other context :", round(ledger.belief(claim, context="segment:new"), 3)) print("all contexts :", round(ledger.belief(claim), 3))unobserved : 0.5 4 confirms : 0.833 4 confirms, 1 contra : 0.556 + PARTIAL 0.5 : 0.5 other context : 0.5 all contexts : 0.5
Four confirmations give (1 + 4) / (5 + 1) = 0.833. One contradiction then counts three times: 5 / (5 + 4) = 0.556. A partial outcome at strength 0.5 adds half a confirmation and half a contradiction. The
segment:newcontext has no observations, so it reads 0.5. The pooled belief equals the returning segment's, also 0.5 here, because every observation so far is in that segment. - Route decisions by confidence
Name each claim after the rule whose decisions it tracks, and use the context for the slice you want to judge separately, here the customer segment. After each decision you learn the outcome of, record it. Before each new decision, ask the ledger how much to trust the rule in that segment. The script also saves and restores the ledger with
snapshot()andrestore(), which ship in the next release; on 0.2.1, see the note in step 3:import json import os from pathlib import Path from helixor_runtime import HelixorBeliefLedger STATE = Path("ledger.json") CLAIM = "refunds.auto_approve_under_50" REVIEW_BELOW = 0.70 def load_ledger() -> HelixorBeliefLedger: """Restore the ledger from its last snapshot, or start empty.""" if STATE.exists(): return HelixorBeliefLedger.restore(json.loads(STATE.read_text())) return HelixorBeliefLedger() def save_ledger(ledger: HelixorBeliefLedger) -> None: """Write a snapshot atomically: a crash leaves the previous file intact.""" tmp = STATE.with_suffix(".tmp") tmp.write_text(json.dumps(ledger.snapshot())) os.replace(tmp, STATE) def route(ledger, claim, context) -> str: """What the application does with the next decision made by this rule.""" if ledger.is_dead(claim, context=context): return "disable_rule" if ledger.is_bedrock(claim, context=context): return "auto" if ledger.belief(claim, context=context) < REVIEW_BELOW: return "auto_with_review" return "auto_sampled" STATE.unlink(missing_ok=True) ledger = load_ledger() # Outcomes reported after the rule auto-approved refunds, per customer segment. history = { "segment:returning": ["CONFIRM"] * 24, "segment:new": ["CONFIRM"] * 6 + ["CONTRADICT"] * 2, "segment:reseller": ["CONTRADICT"] * 4 + ["CONFIRM"], } for context, outcomes in history.items(): for outcome in outcomes: ledger.record(CLAIM, outcome, context=context) print(f"{'context':20} {'belief':>7} bedrock dead route") for context in history: print(f"{context:20} {ledger.belief(CLAIM, context=context):7.3f} " f"{ledger.is_bedrock(CLAIM, context=context)!s:7} " f"{ledger.is_dead(CLAIM, context=context)!s:5} {route(ledger, CLAIM, context)}") # One chargeback in the returning segment. ledger.record(CLAIM, "CONTRADICT", context="segment:returning") print("\nafter one chargeback in segment:returning:") print(f" belief={ledger.belief(CLAIM, context='segment:returning'):.3f} " f"bedrock={ledger.is_bedrock(CLAIM, context='segment:returning')} " f"route={route(ledger, CLAIM, 'segment:returning')}") # Restart: save a snapshot, restore it, and check nothing was lost. save_ledger(ledger) restored = load_ledger() same = all( restored.belief(CLAIM, context=c) == ledger.belief(CLAIM, context=c) for c in history ) print(f"\nsnapshot {json.loads(STATE.read_text())['schema_version']}: beliefs identical: {same}") print(f"evidence in segment:new: {restored.evidence_mass(CLAIM, context='segment:new')}, " f"posterior (alpha, beta): {restored.posterior(CLAIM, context='segment:new')}")python outcome_memory.py
context belief bedrock dead route segment:returning 0.962 True False auto segment:new 0.500 False False auto_with_review segment:reseller 0.133 False True disable_rule after one chargeback in segment:returning: belief=0.862 bedrock=False route=auto_sampled snapshot helixor.belief_ledger.v1: beliefs identical: True evidence in segment:new: 8.0, posterior (alpha, beta): (7.0, 7.0)
Run against the runtime source on 2026-09-28.
Read the table row by row:
- Returning customers: 24 confirmations give 25 / 26 = 0.962 with 24 observations, which clears both bedrock thresholds, so the rule runs without review.
- New customers: 6 confirmations and 2 contradictions give 7 / (7 + 7) = 0.5. A 75% success rate does not earn trust when failures count three times, so each decision is reviewed.
- Resellers: 1 confirmation and 4 contradictions give 2 / (2 + 13) = 0.133 with 5 observations, which is dead: switch the rule off for this segment and send these refunds to a person.
A single chargeback in the returning segment drops belief from 0.962 to 25 / (25 + 4) = 0.862 and removes bedrock status, so that segment moves to sampled review until confirmations rebuild it. This is the asymmetry by design: an established rule loses trust quickly and regains it slowly.
- Persist the ledger
The script above already does this.
save_ledger()writesledger.snapshot()toledger.jsonthrough a temporary file, so a crash mid-write leaves the previous snapshot intact, andload_ledger()restores it at start-up. The last lines of the output confirm the restored ledger gives identical beliefs and the same evidence. A snapshot looks like this:{"schema_version": "helixor.belief_ledger.v1", "pain_multiplier": 3.0, "entries": {"refunds.auto_approve_under_50": {"confirmations": 31.0, "contradictions": 7.0, "contexts": {"segment:returning": [24.0, 1.0], "segment:new": [6.0, 2.0], "segment:reseller": [1.0, 4.0]}}}}Each entry holds the pooled tallies and, per context,
[confirmations, contradictions]. Save after each batch of outcomes, or on a timer, and keep one writer. A snapshot keeps totals, not individual observations; if you want to rebuild with a different multiplier or for a date range later, also keep your own log of outcomes.Known issue in 0.2.1
HelixorBeliefLedgerhas no snapshot or restore, and it does not expose the evidence count or the posterior parameters. On 0.2.1, persist by appending each outcome to a log (one JSON line withclaim,outcome,strengthandcontext) and replaying it into a fresh ledger withrecord()at start-up.Fixed in the next release:
snapshot(),restore(),evidence_mass,posterior,varianceandclaims(), as used above.BeliefSnapshotErroris exported fromhelixor_runtime.
How it works#
Each claim has one tally for all observations and one per context. Recording with a context updates both, so a claim's pooled belief always reflects every segment, while belief(claim, context=...) reflects only that segment. A context you have never recorded against reads 0.5, not the pooled value, so a new segment starts undecided instead of inheriting trust it has not earned.
Next to a decision pack, the usual pattern is:
- Decide with the pack and store the receipt, the rule IDs that fired and the context with your record.
- When the outcome is known, call
record()for each rule involved, with that context. - Before acting on the next decision from a rule, call
route()to choose automatic, sampled, reviewed or stopped handling.
The ledger does not change the pack. When a rule is wrong for a whole segment, the fix is a new pack version, which is what the learning loop proposes and gates.
Limits#
- The ledger lives in one process's memory and has no locking. Guard it with a lock if you share it across threads, and keep one writer for the outcome log.
- Every observation weighs the same regardless of age; there is no time decay. To forget old outcomes, rebuild from the recent part of your own outcome log.
- Belief is an estimate from the outcomes you report. It is only as good as your outcome reporting, and it is not a calibrated probability of the next decision being right.