Evaluating fit
Test the platform on one real decision before you commit. This page gives you a pilot plan that includes the learning loop, a questionnaire to screen candidate decisions, and a worksheet for the cost case. You supply every number.
Pilot plan#
A pilot takes two to four weeks with one engineer, one reviewer who knows the decision, and the decision's business owner. It runs in shadow mode, so it never changes what your users see, and it exercises the whole platform: the core decides, outcomes come back, gaps are classified, and one policy change goes through the gate.
- Pick one decision (days 1 to 3)
Choose a single decision with volume, a definable notion of "correct", outcomes you can observe within the pilot window, and an owner who will act on the result. Write it as facts, invariants and a closed set of actions in the Playbooks format. Use the fit questionnaire below to choose.
- Build a golden set (days 2 to 5)
Collect a few hundred representative historical cases. Label each with the action it should get and the rule that should decide it. Include boundary cases, such as a value exactly on a limit, and cases you expect policy not to settle, so expected escalations are recorded too. This set is the regression gate for every later change. See Testing policies.
- Capture outcomes and overrides from week 1
Wire the feed before shadow mode starts. For every decision, record the receipt and pack version. Then attach what actually happened (loss, dispute, reversal, reassignment, completion) and any override with a reason code. Outcome memory accumulates from these records; without them there is nothing for the learning loop to learn from.
- Run in shadow mode (weeks 1 to 3)
Evaluate live traffic next to your current decision and keep acting on the current one. Log the receipt, action, rule IDs and remedy summary, never the raw input. Compare the two decisions case by case. See Evaluating decisions.
- Measure call time where it matters
Time the call at the caller, in your placement: in process, sidecar or shared service.
latency_uscovers evaluation only. See Performance efficiency. - Classify the gaps
For each case where the decision and the outcome, or the decision and a reviewer, disagree, record the class: data defect (bad or missing input), policy defect (the rule is wrong or missing) or expected (the policy did what it should). Measure the mix. Data defects point to upstream fixes, policy defects to the next step, and a large share of cases that need judgement points to reasoning and human review.
- Run one gated policy change
Take one policy defect. Propose a change, such as an adjusted threshold or a new exception, either from the hosted learning loop Preview or written by your reviewer. Replay the golden set: the change must resolve at least one gap and cause zero regressions. Admit it explicitly as a new pack version; the previous version stays available for rollback. Admission creates a numbered version and keeps the previous ones, so you can roll back.
- Decide (last days)
Fill in the ROI worksheet with what you measured. Decide one of three things: enforce on this decision, extend to a second decision, or stop. Record the reason either way.
What you need for the pilot
The embedded core, outcome memory and the first pack are Available. Compiling your own playbook needs a Developer license, which carries a per-minute decision limit; check it against shadow traffic. In the 0.2.1 compiled pack your own rules are regex and checksum rules; a decision that needs your own codons and invariants, and the hosted learning loop, reasoning and solvers, run in Preview by arrangement. See Licensing and tiers.
Fit questionnaire#
Answer these for each candidate decision. Mostly "yes" in the first group and "no" in the second means it is a good pilot.
Points toward a fit#
| # | Question | Why it matters |
|---|---|---|
| 1 | Can the policy be written down as conditions over facts you can extract from the input? | That is what the core evaluates. |
| 2 | Is there a definable notion of a correct decision? | Without it, neither the gate nor the learning loop has a target. |
| 3 | Can you observe outcomes or overrides, and link them to decisions? | They are what the learning loop learns from. |
| 4 | Does the decision run on every item, at volume, within a latency budget? | The core decides in process, in microseconds, without a model call. |
| 5 | Must the input stay inside your process or perimeter? | The core makes no network calls while it decides. |
| 6 | Will someone ask later why a case was decided as it was, or what would have changed it? | Rules fired, remedies and receipts answer that. |
| 7 | Do several services or languages need the same policy? | One versioned pack replaces copies. |
| 8 | Does the policy need to improve over time without silent changes? | Changes arrive as gated, explicitly admitted versions. |
Points away from a fit#
| # | Question | If yes |
|---|---|---|
| 9 | Is the decision open-ended judgement with no repeatable criteria? | Keep it with people; use reasoning for the evidence-backed part. |
| 10 | Is the core signal perception over raw data, such as images, audio or event sequences? | A trained model is the right core; feed its score to Helixor as a fact. |
| 11 | Must the policy shift minute by minute with no review? | A predictive model fits better; Helixor adapts by gated version. |
| 12 | Must non-engineers change rules daily through a visual tool? | A central rules platform may fit better. |
| 13 | Is volume low, with a person already reviewing every item and no latency pressure? | The case rests on consistency and audit, not cost. |
ROI worksheet#
This worksheet is a method, not a claim. Fill in each input from your own systems and the pilot. Where you cannot measure something, write a range and carry both ends through.
Inputs#
| Symbol | Input | Where to get it |
|---|---|---|
V | Decisions per month | Your traffic or pipeline metrics |
C_now | Current cost per item of the decision as you make it today (tokens, API fees, compute) | Your invoices or cost reports |
T_now | Time the current decision adds per item, at the caller | Your tracing |
T_hx | Time the core adds per decision, measured at the caller in your placement | Pilot step 4 |
B | Latency budget for the decision | Your service-level objective |
R | Measured evaluations per second per process | The benchmark on your hardware |
P_core | Your price per core-hour | Your cloud or data-center cost |
L | License cost per month | Your quote |
e | Share of decisions escalated to hosted reasoning | Your escalation rule, measured in shadow mode |
C_esc | Cost per escalated item | Your hosted pricing |
r_now, r_hx | Share of decisions sent to manual review, today and with Helixor | Your queue metrics; pilot step 5 |
m, W | Minutes per manual review, and loaded cost per reviewer minute | Your operations team |
N_inc | Expected incidents per year from this decision going wrong (a wrong outcome, data reaching the wrong place, a decision you cannot explain) | Your incident history or risk register |
C_inc | Cost per incident: response, notification, remediation | Your security and legal teams |
o_now, o_new | Override rate before and after the gated policy change | Pilot steps 3 and 7 |
m_o | Minutes spent per override | Your operations team |
k | Share of those incidents that the rules in your pack would have decided | Your incident reviews, replayed against the golden set |
Formulas#
| Line | Formula | Notes |
|---|---|---|
| Current decision cost, per month | V × C_now | Zero if the decision is not automated today. |
| Runtime compute, per month | (V ÷ R ÷ 3600) × P_core | Core-hours of pure evaluation. Add headroom for peaks and for serialization if you use the service. |
| Escalation, per month | V × e × C_esc | Only for the share that policy cannot settle. |
| Manual review change, per month | V × (r_now − r_hx) × m × W | Positive is a saving. Can be negative if more cases are referred. |
| Net run cost change, per month | V × C_now − [(V ÷ R ÷ 3600) × P_core + L + V × e × C_esc] + V × (r_now − r_hx) × m × W | Positive means the platform path costs less to run. |
| Override handling saved, per month | V × (o_now − o_new) × m_o × W | What one gated policy change is worth; repeat per admitted change. |
| Latency headroom | B − T_hx versus B − T_now | Whether the decision fits your budget at all is often decisive before cost is. |
| Exposure addressed, per year | N_inc × k × C_inc | An upper bound on what the platform can affect, not a saving. Discount it by the miss rate you measured. |
Keep the worksheet honest
Use the disagreement and miss rates from your own golden set and shadow run, not a published figure. Count one-time costs separately: integration work, pack authoring, the audit store and monitoring. The Cost optimization page explains how to size compute from a measurement.