helixordevelopers

Choosing an approach

There are many ways to build a system that makes decisions. This page sets out Helixor's architecture and compares it with seven common alternatives, criterion by criterion. It also says where each alternative is the better choice.

The Helixor architecture#

Helixor separates three concerns that other approaches tend to merge into one. Deciding happens in a deterministic core. Improving happens in a learning loop around it. Judging what policy cannot settle happens in a reasoning service that is allowed to say "I don't know".

PropertyWhat it meansStatus
Decisions are versioned artifactsA decision policy is declared in a playbook and compiled into a pack with an ID and version. It is deployed alongside your code, not written into it. Every result names the pack version that produced it. PacksAvailable
Typed state, separate from policyExtraction steps (codons) turn raw input into typed facts, and invariants are evaluated over those facts. You can test and improve each layer on its own. ConceptsAvailable
Deterministic, closed outcomesThe same input and pack version always give the same action and receipt. A pack declares every action it can return, so callers can handle every case. Evaluating decisionsAvailable
Counterfactual remediesEach decision carries the smallest change that would alter the outcome, so a refusal comes with what would change it. ConceptsAvailable in the first pack; general counterfactual analysis Preview
A receipt per decisionA fingerprint of pack, action, rules and input that you log against your request. ReceiptsAvailable; signed, chained receipts Planned
Runs next to the dataThe core runs in your process or a sidecar, in microseconds, with no model calls and no network traffic while it decides. BenchmarksAvailable
Outcome memoryRecords what happened after each decision and keeps a confidence per claim and context that weights failures more heavily than successes. It can mark a claim as established or as dead.Available embedded, standalone
Gated learning loopOutcomes and expert overrides are ingested, and each gap is classified as a data defect, a policy defect or expected performance. The loop proposes policy changes and runs them against a golden regression suite. A new pack version is admitted only when it passes. How learning worksPreview
Reasoning that can abstainQuestions policy cannot settle go to a reasoning service. It returns answer, abstain or refuse, with a correctness estimate and a proof reference. It abstains rather than guesses. Hosted escalationPreview
Optimization over the same modelSolvers for allocation, routing and rostering work from the same typed decision state.Preview

How learning works#

Helixor learns around the decision, not inside it. A decision made today can be reproduced exactly next year, because it was made by a fixed pack version. The policy still improves, one reviewed version at a time.

  1. Decide

    The core decides with pack version N and emits a receipt.

  2. Observe

    You report what actually happened, and any expert override, against the decision. Outcome memory updates its confidence for the claims involved.

  3. Diagnose

    Each gap between decision and outcome is classified. A data defect is fixed at the source, expected performance is left alone, and a policy defect is a candidate for change.

  4. Propose

    The learning loop generates candidate changes, such as adjusted thresholds or new exceptions, that would have resolved the policy defects.

  5. Gate

    Each candidate is replayed against a golden set of past decisions. It must resolve at least one gap and cause zero regressions.

  6. Admit

    A passing candidate runs only after it is explicitly admitted: a preflight simulation, then a new numbered pack version that becomes active. The reviewer must be a different identity from the one that proposed it; self-review is refused. Every version and an audit log stay available, and you can roll back to any earlier version. See Close the learning loop.

This is the difference from a model that retrains itself: nothing changes in production without passing the gate and being explicitly admitted, and every change is a numbered version you can inspect and roll back. It is also the difference from a static rules engine: outcomes drive the proposals, so policies do not depend on someone noticing a problem.

The approaches compared#

Each column is a class of tools, not a product, and describes the typical case. Many real systems combine several.

ApproachWhat it means here
HelixorThe architecture above: deterministic core, gated learning loop, reasoning that can abstain.
Language model or agentA language model, or an agent built on one, decides from a prompt and context.
Predictive modelA model trained on historical outcomes returns a score or class.
Central rules platformA business-rules or decision-management platform that applications call over the network.
Policy-as-code enginePolicies in a policy language, evaluated by a general engine; often used for authorization.
Hand-coded logicConditions written directly in each application service.
Human reviewPeople decide each case from a queue.
CriterionHelixorLanguage model / agentPredictive modelCentral rules platformPolicy-as-codeHand-codedHuman review
RepeatabilityExact, per pack versionCan vary between calls and versionsFixed per model versionExact per rule versionExactExactVaries by reviewer
ExplanationRules fired, reason, remedy, receipt on every decisionA generated rationale, not the mechanism that decidedAttributions, if addedDecision logs, rule historyPolicy and input, if loggedWhatever you logNotes, if recorded
Improves from outcomesYes, through gated versions (preview)Through prompt, retrieval or fine-tuning changesYes, by retraining; drift must be monitoredManually, by authorsManuallyManuallyPeople learn, unevenly
Change controlEvery change regression-gated, admitted by a second person, versioned and reversibleHard to test exhaustivelyRetrain, validate, redeployStrong authoring and approvalsSource control and reviewCode review per serviceTraining and guidance
Ambiguity and unstructured meaningEscalates to reasoning; abstains when unsureStrong, but does not know when it is wrongOnly what it was trained onNeeds structured factsNeeds structured factsWeakStrong
Knows when it cannot decideYes: closed actions, abstention, fail-closed errorsRarely; confident errors are possibleNo; scores outside its training range look normalDefault rule or errorUsually default-denyWhatever the code doesYes, if escalation exists
Latency and placementMicroseconds in process; hosted parts only when neededA model call, usually remoteLocal or remote scoringA network round tripLibrary or sidecarIn processMinutes to days
Data egress to decideNone for the core; only what you escalatePrompt and context, unless self-hostedFeatures, if remoteThe decision inputNone if localNoneStays with staff
Cost per decisionCPU microseconds; reasoning only for the remainderTokens or accelerator time on every callInference computePer-call or capacity pricingCPUCPUStaff time
OptimizationSolvers over the same decision state (preview)Not reliable for constrained optimizationNoUsually separate productsNoCustom codeManual

Where each alternative is the better choice#

  • Language model or agent: open-ended tasks with no definable "correct": drafting, summarizing, exploring. Use Helixor around it to guard inputs and outputs and to decide what it may act on.
  • Predictive model: perception and pattern recognition over raw signals, such as images, audio or behavioral sequences, where features cannot be written down. A model's score can be one input to a Helixor decision.
  • Central rules platform: when business users must author rules through a visual tool and every caller can tolerate a network hop.
  • Policy-as-code engine: infrastructure authorization (who may call what), where a mature ecosystem already exists.
  • Hand-coded logic: a single simple condition in one service that nobody will audit.
  • Human review: rare, high-stakes cases, and anything the reasoning service abstains on.

What the architecture asks of you#

  • A definition of correct. Helixor decides what can be stated as invariants and objectives, and learns from outcomes you report. If nobody can say what a good decision is, no learning loop can find it; the platform routes those cases to reasoning or people instead of guessing.
  • Investment up front. A pack needs a playbook, extraction for the facts it uses, and a golden set of cases. That work is what makes decisions testable, explainable and safe to evolve.
  • Deliberate, not instant, adaptation. Policies improve version by version through the gate, not continuously. That is slower than online learning by design; if a decision must shift minute by minute with no review, a predictive model is the better core.
  • Outcome reporting. The learning loop improves only as fast as you report outcomes and overrides.

Current release

The embedded runtime is at 0.2.1. It is Python-first, with other languages served through the decision service. Compiled packs currently run the built-in checks plus regex and checksum rules. Outcome memory ships embedded; the learning loop, reasoning and solvers run on the hosted Helixor platform and are in Preview. The User Guide documents what each release supports exactly.