Benchmarks
Measured results for Helixor components, with the hardware they ran on, every raw run, and the scripts to reproduce them. We publish absolute numbers only, and only numbers you can reproduce.
Results
Decision Runtime
About 25 µs per in-process decision at the median, streaming per push, cold start and HTTP round trip.
MethodMethodology
What each metric covers, how runs are set up, and the reporting rules.
MethodReproduce the results
Download the scripts and measure on your own hardware.
Headline results#
| Component | Result | Conditions |
|---|---|---|
| Decision Runtime, in-process | 24.9 µs p50, 48.8 µs p99 per evaluate() call | Runtime 0.2.1 (Python build), built-in pack, Apple M5 Max, single process, load average 7.5 to 10.5 |
| Decision Runtime, sidecar | 347 µs p50 per HTTP round trip on loopback | Same machine, one keep-alive connection |
Coming next#
| Area | What will be published | Status |
|---|---|---|
| Decision Runtime on server hardware | Linux x86_64 and aarch64, multi-process scaling, memory per engine, payload-size sensitivity | Planned |
| Native Decision Runtime | The same suite against the compiled runtime | Planned |
| Rostering solver | Time to a feasible roster on public rostering competition instances, with the dataset and scripts | Planned |
| Hosted reasoning | Answer, abstain and refuse rates with calibration on a public, held-out question set | Planned |
A result appears here only when its scripts, data and environment can be published with it. See Methodology.