Decision Runtime results
Measured latency and throughput of runtime 0.2.1 on one machine, three runs per benchmark. Every figure links to the script that produced it and to the raw output of all runs.
Read these before the numbers
- Python build. These results measure the current pure-Python runtime. A compiled native runtime is Planned; its results will be published separately when it ships.
- Not an idle machine. Load average was 7.5 to 10.5 during the runs, so these are conservative rather than best-case, and tail values (p99, max) are noisier than they would be on a quiet host. The 0.2.0 runs had a lower load (5.6 to 7.1), so small differences between the two releases are within noise.
- One laptop. Use these for order of magnitude and shape. Measure on your own hardware with the published scripts.
Environment#
| CPU | Apple M5 Max, 18 cores (6 performance, 12 efficiency) |
| Memory | 48 GB |
| OS | macOS 26.6.2 |
| Software | Python 3.11.4, runtime 0.2.1, pydantic 2.13.5; HTTP: fastapi 0.123.0, uvicorn 0.40.0 |
| Workload | Five fixed payloads cycled in order: one clean, one SSN + IP, one card + email, one health record ID, one phone + email. Built-in Community pack. Single process. |
Summary#
Medians of three runs. Metric definitions are on Methodology.
| Measurement | p50 | p90 | p99 | Throughput |
|---|---|---|---|---|
Engine time (latency_us) | 19.6 µs | 23.9 µs | 59.6 µs | 36,048 evaluations/s (loop) |
Call time, evaluate() end to end | 24.9 µs | 32.3 µs | 48.8 µs | 36,495 calls/s (1 / mean) |
Streaming, one push() of 8 characters | 57.9 µs | 75.8 µs | 159.4 µs | 2 KB stream in 15.1 ms |
HTTP round trip, POST /v1/evaluate on loopback | 347 µs | 414 µs | 648 µs | 2,739 requests/s (one connection, sequential) |
Cold start#
| Stage | Fresh process (5 runs) |
|---|---|
Import helixor_runtime | 864 to 985 ms |
Construct HelixorEngine() | 0.21 to 0.24 ms |
First evaluate() | 107.5 to 112.6 µs |
Second evaluate() | 48.0 to 55.1 µs |
Cold start in the next release#
Re-measured on 2026-09-28 with the runtime built from source after the fix that stops importing the GPU solver's libraries when you import the runtime. No release has been cut; __version__ still reads 0.2.1. Same machine, load averages 3.25, 4.19 and 5.91 (1, 5 and 15 minutes); only cold start was re-run.
| Stage | Fresh process (5 runs) |
|---|---|
Import helixor_runtime | 107.0 to 108.6 ms |
Construct HelixorEngine() | 0.38 to 0.42 ms |
First evaluate() | 91.0 to 104.0 µs |
Second evaluate() | 44.7 to 53.3 µs |
| Whole process (start, import, construct, two evaluations) | 0.15 s wall time in every run |
Cold start target not met
The target is a cold start under 100 ms. Import alone takes about 108 ms, so the target is not met in 0.2.1 or on main. This is an open issue.
Raw results: results-2026-09-28-cold-start.md.
What the numbers mean#
- About 25 µs per in-process decision at the median, including license check, receipt hashing and building the result. Budget from this figure, not from
latency_us, which is about 5 µs lower because it covers evaluation only. - A loopback HTTP hop costs about 14 times the decision. The sidecar pattern adds roughly 0.3 ms; the evaluation inside it is still about 28 µs. Put the runtime in-process when the caller is Python and latency matters. See Performance.
- Each streaming push costs about three evaluations. The session re-checks its look-ahead buffer on every push, and before it releases text it evaluates the parts on each side of the cut, so that no value is split. That check is what fixed the damaged text of 0.2.0, and it makes a push about three times the 0.2.0 cost. Pushing larger fragments means fewer checks.
- Import time dominates cold start. In 0.2.1, about 0.6 s of the 0.9 s is the GPU solver module and the optional numeric library it imports (measured with
python -X importtime), which the example pack does not use. In the next release that import is lazy and importing the runtime takes about 108 ms, most of it the engine's own modules and the validation library (see above). That is still above the 100 ms target. For Lambda-style platforms, import at initialization, not per request. - Warm up once. The first evaluation in a process is about twice as slow as the second, and several times slower than steady state. Run one clean evaluation at startup.
All runs#
Engine time: bench_engine_time.py#
These runs used the earlier example script, which measured the same thing with the same payloads and count; bench_engine_time.py reproduces it.
| Run | Evaluations/s | Mean | p50 | p90 | p99 |
|---|---|---|---|---|---|
| 1 | 36,048 | 20.77 µs | 19.62 µs | 23.92 µs | 59.58 µs |
| 2 | 20,774 | 34.68 µs | 25.58 µs | 63.58 µs | 135.92 µs |
| 3 | 41,277 | 18.38 µs | 18.58 µs | 21.29 µs | 27.42 µs |
Call time: bench_evaluate_e2e.py#
| Run | p50 | p90 | p99 | Max | Calls/s |
|---|---|---|---|---|---|
| 1 | 24.12 µs | 26.75 µs | 33.96 µs | 182.67 µs | 41,670 |
| 2 | 24.88 µs | 32.29 µs | 48.83 µs | 6,215.04 µs | 36,495 |
| 3 | 26.88 µs | 33.00 µs | 75.00 µs | 267.38 µs | 34,914 |
Streaming: bench_streaming.py#
| Run | push p50 | p90 | p99 | Max | 2 KB stream p50 |
|---|---|---|---|---|---|
| 1 | 57.92 µs | 75.75 µs | 159.42 µs | 307.62 µs | 15.11 ms |
| 2 | 59.46 µs | 76.96 µs | 159.67 µs | 280.04 µs | 15.89 ms |
| 3 | 53.83 µs | 66.46 µs | 130.88 µs | 175.75 µs | 14.06 ms |
HTTP: bench_http.py#
| Run | p50 | p90 | p99 | Max | Requests/s | Server engine p50 |
|---|---|---|---|---|---|---|
| 1 | 340 µs | 369 µs | 488 µs | 1,282 µs | 2,889 | 27.5 µs |
| 2 | 347 µs | 414 µs | 687 µs | 1,559 µs | 2,739 | 28.0 µs |
| 3 | 360 µs | 416 µs | 648 µs | 1,222 µs | 2,670 | 29.2 µs |
Raw results: results-2026-09-28.md.
History#
Earlier results, kept as published:
- results-2026-09-27.md: runtime 0.2.0, load average 5.6 to 7.1. Medians: call time 22.3 µs p50, streaming push 19.2 µs p50 (before the 0.2.1 streaming fix), HTTP round trip 339 µs p50, import 684 to 728 ms.
Not measured yet#
- Multi-process throughput scaling on a server CPU.
- Payload-length sensitivity (cost against text size and match count).
- Memory footprint per engine.
- Concurrent HTTP clients against one service.
- Linux on x86_64 and aarch64 server hardware.
These are next; they will be added here with their scripts.