Reproduce the results
Every published number comes from one of the scripts on this page. Run them on the hardware you will deploy to; your numbers are the ones that matter.
Prerequisites#
- The runtime installed in a virtual environment (Installation). Python 3.11 matches the published runs.
- The scripts below, saved together in one directory (they share
_common.py). - An otherwise idle machine on mains power.
Availability
The scripts only need the helixor-runtime package. Until it is on the public package index, use the wheel from your account.
Scripts#
| Script | Measures |
|---|---|
_common.py | Shared payloads, environment report and percentile helpers. |
bench_engine_time.py | Engine time: the engine-reported latency_us over 5,000 evaluations. |
bench_evaluate_e2e.py | Call time: 100 warm-up then 10,000 timed evaluate() calls. |
bench_cold_call.py | Import, construction, first and second evaluation in a fresh process. |
bench_streaming.py | Per-push() time streaming a 2 KB text in 8-character fragments. |
bench_http.py | Round trip of POST /v1/evaluate over loopback, one keep-alive connection. |
Run#
cd bench/ # engine time, call time, streaming and HTTP, three runs each for s in bench_engine_time bench_evaluate_e2e bench_streaming bench_http; do for i in 1 2 3; do python $s.py; done done # cold start: five fresh processes for i in 1 2 3 4 5; do python bench_cold_call.py; done
Each script prints the environment first, then one summary line per measurement:
cpu=Apple M5 Max cores=18 ram=48 GB os=macOS-26.6.2-arm64-arm-64bit python=3.11.4 evaluate() end-to-end wall clock: n=10000 p50=24.12us p90=26.75us p99=33.96us max=182.67us mean=24.00us ops/s=41,670 evaluate() engine-reported latency_us: n=10000 p50=18.50us p90=20.96us p99=26.00us max=144.08us mean=18.23us ops/s=54,840
macOS: numeric library conflict
If the import aborts with an OpenMP "duplicate library" message, run in a clean virtual environment, or set KMP_DUPLICATE_LIB_OK=TRUE for the benchmark run.
Compare and share#
- Compare medians of three runs, not single runs. Expect p99 and maximum to vary most.
- Measure your own payloads too: replace
PAYLOADSin_common.pywith a sample of real (test) traffic. - If your results differ a lot from ours on similar hardware, email hello@helixor.ai with the full script output.