helixordevelopers

Evaluating fit

Test the platform on one real decision before you commit. This page gives you a pilot plan that includes the learning loop, a questionnaire to screen candidate decisions, and a worksheet for the cost case. You supply every number.

Pilot plan#

A pilot takes two to four weeks with one engineer, one reviewer who knows the decision, and the decision's business owner. It runs in shadow mode, so it never changes what your users see, and it exercises the whole platform: the core decides, outcomes come back, gaps are classified, and one policy change goes through the gate.

  1. Pick one decision (days 1 to 3)

    Choose a single decision with volume, a definable notion of "correct", outcomes you can observe within the pilot window, and an owner who will act on the result. Write it as facts, invariants and a closed set of actions in the Playbooks format. Use the fit questionnaire below to choose.

  2. Build a golden set (days 2 to 5)

    Collect a few hundred representative historical cases. Label each with the action it should get and the rule that should decide it. Include boundary cases, such as a value exactly on a limit, and cases you expect policy not to settle, so expected escalations are recorded too. This set is the regression gate for every later change. See Testing policies.

  3. Capture outcomes and overrides from week 1

    Wire the feed before shadow mode starts. For every decision, record the receipt and pack version. Then attach what actually happened (loss, dispute, reversal, reassignment, completion) and any override with a reason code. Outcome memory accumulates from these records; without them there is nothing for the learning loop to learn from.

  4. Run in shadow mode (weeks 1 to 3)

    Evaluate live traffic next to your current decision and keep acting on the current one. Log the receipt, action, rule IDs and remedy summary, never the raw input. Compare the two decisions case by case. See Evaluating decisions.

  5. Measure call time where it matters

    Time the call at the caller, in your placement: in process, sidecar or shared service. latency_us covers evaluation only. See Performance efficiency.

  6. Classify the gaps

    For each case where the decision and the outcome, or the decision and a reviewer, disagree, record the class: data defect (bad or missing input), policy defect (the rule is wrong or missing) or expected (the policy did what it should). Measure the mix. Data defects point to upstream fixes, policy defects to the next step, and a large share of cases that need judgement points to reasoning and human review.

  7. Run one gated policy change

    Take one policy defect. Propose a change, such as an adjusted threshold or a new exception, either from the hosted learning loop Preview or written by your reviewer. Replay the golden set: the change must resolve at least one gap and cause zero regressions. Admit it explicitly as a new pack version; the previous version stays available for rollback. Admission creates a numbered version and keeps the previous ones, so you can roll back.

  8. Decide (last days)

    Fill in the ROI worksheet with what you measured. Decide one of three things: enforce on this decision, extend to a second decision, or stop. Record the reason either way.

What you need for the pilot

The embedded core, outcome memory and the first pack are Available. Compiling your own playbook needs a Developer license, which carries a per-minute decision limit; check it against shadow traffic. In the 0.2.1 compiled pack your own rules are regex and checksum rules; a decision that needs your own codons and invariants, and the hosted learning loop, reasoning and solvers, run in Preview by arrangement. See Licensing and tiers.

Fit questionnaire#

Answer these for each candidate decision. Mostly "yes" in the first group and "no" in the second means it is a good pilot.

Points toward a fit#

#QuestionWhy it matters
1Can the policy be written down as conditions over facts you can extract from the input?That is what the core evaluates.
2Is there a definable notion of a correct decision?Without it, neither the gate nor the learning loop has a target.
3Can you observe outcomes or overrides, and link them to decisions?They are what the learning loop learns from.
4Does the decision run on every item, at volume, within a latency budget?The core decides in process, in microseconds, without a model call.
5Must the input stay inside your process or perimeter?The core makes no network calls while it decides.
6Will someone ask later why a case was decided as it was, or what would have changed it?Rules fired, remedies and receipts answer that.
7Do several services or languages need the same policy?One versioned pack replaces copies.
8Does the policy need to improve over time without silent changes?Changes arrive as gated, explicitly admitted versions.

Points away from a fit#

#QuestionIf yes
9Is the decision open-ended judgement with no repeatable criteria?Keep it with people; use reasoning for the evidence-backed part.
10Is the core signal perception over raw data, such as images, audio or event sequences?A trained model is the right core; feed its score to Helixor as a fact.
11Must the policy shift minute by minute with no review?A predictive model fits better; Helixor adapts by gated version.
12Must non-engineers change rules daily through a visual tool?A central rules platform may fit better.
13Is volume low, with a person already reviewing every item and no latency pressure?The case rests on consistency and audit, not cost.

ROI worksheet#

This worksheet is a method, not a claim. Fill in each input from your own systems and the pilot. Where you cannot measure something, write a range and carry both ends through.

Inputs#

SymbolInputWhere to get it
VDecisions per monthYour traffic or pipeline metrics
C_nowCurrent cost per item of the decision as you make it today (tokens, API fees, compute)Your invoices or cost reports
T_nowTime the current decision adds per item, at the callerYour tracing
T_hxTime the core adds per decision, measured at the caller in your placementPilot step 4
BLatency budget for the decisionYour service-level objective
RMeasured evaluations per second per processThe benchmark on your hardware
P_coreYour price per core-hourYour cloud or data-center cost
LLicense cost per monthYour quote
eShare of decisions escalated to hosted reasoningYour escalation rule, measured in shadow mode
C_escCost per escalated itemYour hosted pricing
r_now, r_hxShare of decisions sent to manual review, today and with HelixorYour queue metrics; pilot step 5
m, WMinutes per manual review, and loaded cost per reviewer minuteYour operations team
N_incExpected incidents per year from this decision going wrong (a wrong outcome, data reaching the wrong place, a decision you cannot explain)Your incident history or risk register
C_incCost per incident: response, notification, remediationYour security and legal teams
o_now, o_newOverride rate before and after the gated policy changePilot steps 3 and 7
m_oMinutes spent per overrideYour operations team
kShare of those incidents that the rules in your pack would have decidedYour incident reviews, replayed against the golden set

Formulas#

LineFormulaNotes
Current decision cost, per monthV × C_nowZero if the decision is not automated today.
Runtime compute, per month(V ÷ R ÷ 3600) × P_coreCore-hours of pure evaluation. Add headroom for peaks and for serialization if you use the service.
Escalation, per monthV × e × C_escOnly for the share that policy cannot settle.
Manual review change, per monthV × (r_now − r_hx) × m × WPositive is a saving. Can be negative if more cases are referred.
Net run cost change, per monthV × C_now − [(V ÷ R ÷ 3600) × P_core + L + V × e × C_esc] + V × (r_now − r_hx) × m × WPositive means the platform path costs less to run.
Override handling saved, per monthV × (o_now − o_new) × m_o × WWhat one gated policy change is worth; repeat per admitted change.
Latency headroomB − T_hx versus B − T_nowWhether the decision fits your budget at all is often decisive before cost is.
Exposure addressed, per yearN_inc × k × C_incAn upper bound on what the platform can affect, not a saving. Discount it by the miss rate you measured.

Keep the worksheet honest

Use the disagreement and miss rates from your own golden set and shadow run, not a published figure. Count one-time costs separately: integration work, pack authoring, the audit store and monitoring. The Cost optimization page explains how to size compute from a measurement.