helixordevelopers

Decision Runtime results

Measured latency and throughput of runtime 0.2.1 on one machine, three runs per benchmark. Every figure links to the script that produced it and to the raw output of all runs.

Run 2026-09-28Runtime 0.2.1Python 3.11.4

Read these before the numbers

  • Python build. These results measure the current pure-Python runtime. A compiled native runtime is Planned; its results will be published separately when it ships.
  • Not an idle machine. Load average was 7.5 to 10.5 during the runs, so these are conservative rather than best-case, and tail values (p99, max) are noisier than they would be on a quiet host. The 0.2.0 runs had a lower load (5.6 to 7.1), so small differences between the two releases are within noise.
  • One laptop. Use these for order of magnitude and shape. Measure on your own hardware with the published scripts.

Environment#

CPUApple M5 Max, 18 cores (6 performance, 12 efficiency)
Memory48 GB
OSmacOS 26.6.2
SoftwarePython 3.11.4, runtime 0.2.1, pydantic 2.13.5; HTTP: fastapi 0.123.0, uvicorn 0.40.0
WorkloadFive fixed payloads cycled in order: one clean, one SSN + IP, one card + email, one health record ID, one phone + email. Built-in Community pack. Single process.

Summary#

Medians of three runs. Metric definitions are on Methodology.

Measurementp50p90p99Throughput
Engine time (latency_us)19.6 µs23.9 µs59.6 µs36,048 evaluations/s (loop)
Call time, evaluate() end to end24.9 µs32.3 µs48.8 µs36,495 calls/s (1 / mean)
Streaming, one push() of 8 characters57.9 µs75.8 µs159.4 µs2 KB stream in 15.1 ms
HTTP round trip, POST /v1/evaluate on loopback347 µs414 µs648 µs2,739 requests/s (one connection, sequential)

Cold start#

StageFresh process (5 runs)
Import helixor_runtime864 to 985 ms
Construct HelixorEngine()0.21 to 0.24 ms
First evaluate()107.5 to 112.6 µs
Second evaluate()48.0 to 55.1 µs

Cold start in the next release#

Re-measured on 2026-09-28 with the runtime built from source after the fix that stops importing the GPU solver's libraries when you import the runtime. No release has been cut; __version__ still reads 0.2.1. Same machine, load averages 3.25, 4.19 and 5.91 (1, 5 and 15 minutes); only cold start was re-run.

StageFresh process (5 runs)
Import helixor_runtime107.0 to 108.6 ms
Construct HelixorEngine()0.38 to 0.42 ms
First evaluate()91.0 to 104.0 µs
Second evaluate()44.7 to 53.3 µs
Whole process (start, import, construct, two evaluations)0.15 s wall time in every run

Cold start target not met

The target is a cold start under 100 ms. Import alone takes about 108 ms, so the target is not met in 0.2.1 or on main. This is an open issue.

Raw results: results-2026-09-28-cold-start.md.

What the numbers mean#

  • About 25 µs per in-process decision at the median, including license check, receipt hashing and building the result. Budget from this figure, not from latency_us, which is about 5 µs lower because it covers evaluation only.
  • A loopback HTTP hop costs about 14 times the decision. The sidecar pattern adds roughly 0.3 ms; the evaluation inside it is still about 28 µs. Put the runtime in-process when the caller is Python and latency matters. See Performance.
  • Each streaming push costs about three evaluations. The session re-checks its look-ahead buffer on every push, and before it releases text it evaluates the parts on each side of the cut, so that no value is split. That check is what fixed the damaged text of 0.2.0, and it makes a push about three times the 0.2.0 cost. Pushing larger fragments means fewer checks.
  • Import time dominates cold start. In 0.2.1, about 0.6 s of the 0.9 s is the GPU solver module and the optional numeric library it imports (measured with python -X importtime), which the example pack does not use. In the next release that import is lazy and importing the runtime takes about 108 ms, most of it the engine's own modules and the validation library (see above). That is still above the 100 ms target. For Lambda-style platforms, import at initialization, not per request.
  • Warm up once. The first evaluation in a process is about twice as slow as the second, and several times slower than steady state. Run one clean evaluation at startup.

All runs#

Engine time: bench_engine_time.py#

These runs used the earlier example script, which measured the same thing with the same payloads and count; bench_engine_time.py reproduces it.

RunEvaluations/sMeanp50p90p99
136,04820.77 µs19.62 µs23.92 µs59.58 µs
220,77434.68 µs25.58 µs63.58 µs135.92 µs
341,27718.38 µs18.58 µs21.29 µs27.42 µs

Call time: bench_evaluate_e2e.py#

Runp50p90p99MaxCalls/s
124.12 µs26.75 µs33.96 µs182.67 µs41,670
224.88 µs32.29 µs48.83 µs6,215.04 µs36,495
326.88 µs33.00 µs75.00 µs267.38 µs34,914

Streaming: bench_streaming.py#

Runpush p50p90p99Max2 KB stream p50
157.92 µs75.75 µs159.42 µs307.62 µs15.11 ms
259.46 µs76.96 µs159.67 µs280.04 µs15.89 ms
353.83 µs66.46 µs130.88 µs175.75 µs14.06 ms

HTTP: bench_http.py#

Runp50p90p99MaxRequests/sServer engine p50
1340 µs369 µs488 µs1,282 µs2,88927.5 µs
2347 µs414 µs687 µs1,559 µs2,73928.0 µs
3360 µs416 µs648 µs1,222 µs2,67029.2 µs

Raw results: results-2026-09-28.md.

History#

Earlier results, kept as published:

  • results-2026-09-27.md: runtime 0.2.0, load average 5.6 to 7.1. Medians: call time 22.3 µs p50, streaming push 19.2 µs p50 (before the 0.2.1 streaming fix), HTTP round trip 339 µs p50, import 684 to 728 ms.

Not measured yet#

  • Multi-process throughput scaling on a server CPU.
  • Payload-length sensitivity (cost against text size and match count).
  • Memory footprint per engine.
  • Concurrent HTTP clients against one service.
  • Linux on x86_64 and aarch64 server hardware.

These are next; they will be added here with their scripts.