Choosing an approach
There are many ways to build a system that makes decisions. This page sets out Helixor's architecture and compares it with seven common alternatives, criterion by criterion. It also says where each alternative is the better choice.
The Helixor architecture#
Helixor separates three concerns that other approaches tend to merge into one. Deciding happens in a deterministic core. Improving happens in a learning loop around it. Judging what policy cannot settle happens in a reasoning service that is allowed to say "I don't know".
| Property | What it means | Status |
|---|---|---|
| Decisions are versioned artifacts | A decision policy is declared in a playbook and compiled into a pack with an ID and version. It is deployed alongside your code, not written into it. Every result names the pack version that produced it. Packs | Available |
| Typed state, separate from policy | Extraction steps (codons) turn raw input into typed facts, and invariants are evaluated over those facts. You can test and improve each layer on its own. Concepts | Available |
| Deterministic, closed outcomes | The same input and pack version always give the same action and receipt. A pack declares every action it can return, so callers can handle every case. Evaluating decisions | Available |
| Counterfactual remedies | Each decision carries the smallest change that would alter the outcome, so a refusal comes with what would change it. Concepts | Available in the first pack; general counterfactual analysis Preview |
| A receipt per decision | A fingerprint of pack, action, rules and input that you log against your request. Receipts | Available; signed, chained receipts Planned |
| Runs next to the data | The core runs in your process or a sidecar, in microseconds, with no model calls and no network traffic while it decides. Benchmarks | Available |
| Outcome memory | Records what happened after each decision and keeps a confidence per claim and context that weights failures more heavily than successes. It can mark a claim as established or as dead. | Available embedded, standalone |
| Gated learning loop | Outcomes and expert overrides are ingested, and each gap is classified as a data defect, a policy defect or expected performance. The loop proposes policy changes and runs them against a golden regression suite. A new pack version is admitted only when it passes. How learning works | Preview |
| Reasoning that can abstain | Questions policy cannot settle go to a reasoning service. It returns answer, abstain or refuse, with a correctness estimate and a proof reference. It abstains rather than guesses. Hosted escalation | Preview |
| Optimization over the same model | Solvers for allocation, routing and rostering work from the same typed decision state. | Preview |
How learning works#
Helixor learns around the decision, not inside it. A decision made today can be reproduced exactly next year, because it was made by a fixed pack version. The policy still improves, one reviewed version at a time.
- Decide
The core decides with pack version N and emits a receipt.
- Observe
You report what actually happened, and any expert override, against the decision. Outcome memory updates its confidence for the claims involved.
- Diagnose
Each gap between decision and outcome is classified. A data defect is fixed at the source, expected performance is left alone, and a policy defect is a candidate for change.
- Propose
The learning loop generates candidate changes, such as adjusted thresholds or new exceptions, that would have resolved the policy defects.
- Gate
Each candidate is replayed against a golden set of past decisions. It must resolve at least one gap and cause zero regressions.
- Admit
A passing candidate runs only after it is explicitly admitted: a preflight simulation, then a new numbered pack version that becomes active. The reviewer must be a different identity from the one that proposed it; self-review is refused. Every version and an audit log stay available, and you can roll back to any earlier version. See Close the learning loop.
This is the difference from a model that retrains itself: nothing changes in production without passing the gate and being explicitly admitted, and every change is a numbered version you can inspect and roll back. It is also the difference from a static rules engine: outcomes drive the proposals, so policies do not depend on someone noticing a problem.
The approaches compared#
Each column is a class of tools, not a product, and describes the typical case. Many real systems combine several.
| Approach | What it means here |
|---|---|
| Helixor | The architecture above: deterministic core, gated learning loop, reasoning that can abstain. |
| Language model or agent | A language model, or an agent built on one, decides from a prompt and context. |
| Predictive model | A model trained on historical outcomes returns a score or class. |
| Central rules platform | A business-rules or decision-management platform that applications call over the network. |
| Policy-as-code engine | Policies in a policy language, evaluated by a general engine; often used for authorization. |
| Hand-coded logic | Conditions written directly in each application service. |
| Human review | People decide each case from a queue. |
| Criterion | Helixor | Language model / agent | Predictive model | Central rules platform | Policy-as-code | Hand-coded | Human review |
|---|---|---|---|---|---|---|---|
| Repeatability | Exact, per pack version | Can vary between calls and versions | Fixed per model version | Exact per rule version | Exact | Exact | Varies by reviewer |
| Explanation | Rules fired, reason, remedy, receipt on every decision | A generated rationale, not the mechanism that decided | Attributions, if added | Decision logs, rule history | Policy and input, if logged | Whatever you log | Notes, if recorded |
| Improves from outcomes | Yes, through gated versions (preview) | Through prompt, retrieval or fine-tuning changes | Yes, by retraining; drift must be monitored | Manually, by authors | Manually | Manually | People learn, unevenly |
| Change control | Every change regression-gated, admitted by a second person, versioned and reversible | Hard to test exhaustively | Retrain, validate, redeploy | Strong authoring and approvals | Source control and review | Code review per service | Training and guidance |
| Ambiguity and unstructured meaning | Escalates to reasoning; abstains when unsure | Strong, but does not know when it is wrong | Only what it was trained on | Needs structured facts | Needs structured facts | Weak | Strong |
| Knows when it cannot decide | Yes: closed actions, abstention, fail-closed errors | Rarely; confident errors are possible | No; scores outside its training range look normal | Default rule or error | Usually default-deny | Whatever the code does | Yes, if escalation exists |
| Latency and placement | Microseconds in process; hosted parts only when needed | A model call, usually remote | Local or remote scoring | A network round trip | Library or sidecar | In process | Minutes to days |
| Data egress to decide | None for the core; only what you escalate | Prompt and context, unless self-hosted | Features, if remote | The decision input | None if local | None | Stays with staff |
| Cost per decision | CPU microseconds; reasoning only for the remainder | Tokens or accelerator time on every call | Inference compute | Per-call or capacity pricing | CPU | CPU | Staff time |
| Optimization | Solvers over the same decision state (preview) | Not reliable for constrained optimization | No | Usually separate products | No | Custom code | Manual |
Where each alternative is the better choice#
- Language model or agent: open-ended tasks with no definable "correct": drafting, summarizing, exploring. Use Helixor around it to guard inputs and outputs and to decide what it may act on.
- Predictive model: perception and pattern recognition over raw signals, such as images, audio or behavioral sequences, where features cannot be written down. A model's score can be one input to a Helixor decision.
- Central rules platform: when business users must author rules through a visual tool and every caller can tolerate a network hop.
- Policy-as-code engine: infrastructure authorization (who may call what), where a mature ecosystem already exists.
- Hand-coded logic: a single simple condition in one service that nobody will audit.
- Human review: rare, high-stakes cases, and anything the reasoning service abstains on.
What the architecture asks of you#
- A definition of correct. Helixor decides what can be stated as invariants and objectives, and learns from outcomes you report. If nobody can say what a good decision is, no learning loop can find it; the platform routes those cases to reasoning or people instead of guessing.
- Investment up front. A pack needs a playbook, extraction for the facts it uses, and a golden set of cases. That work is what makes decisions testable, explainable and safe to evolve.
- Deliberate, not instant, adaptation. Policies improve version by version through the gate, not continuously. That is slower than online learning by design; if a decision must shift minute by minute with no review, a predictive model is the better core.
- Outcome reporting. The learning loop improves only as fast as you report outcomes and overrides.
Current release
The embedded runtime is at 0.2.1. It is Python-first, with other languages served through the decision service. Compiled packs currently run the built-in checks plus regex and checksum rules. Outcome memory ships embedded; the learning loop, reasoning and solvers run on the hosted Helixor platform and are in Preview. The User Guide documents what each release supports exactly.