helixordevelopers

Reproduce the results

Every published number comes from one of the scripts on this page. Run them on the hardware you will deploy to; your numbers are the ones that matter.

Prerequisites#

  • The runtime installed in a virtual environment (Installation). Python 3.11 matches the published runs.
  • The scripts below, saved together in one directory (they share _common.py).
  • An otherwise idle machine on mains power.

Availability

The scripts only need the helixor-runtime package. Until it is on the public package index, use the wheel from your account.

Scripts#

ScriptMeasures
_common.pyShared payloads, environment report and percentile helpers.
bench_engine_time.pyEngine time: the engine-reported latency_us over 5,000 evaluations.
bench_evaluate_e2e.pyCall time: 100 warm-up then 10,000 timed evaluate() calls.
bench_cold_call.pyImport, construction, first and second evaluation in a fresh process.
bench_streaming.pyPer-push() time streaming a 2 KB text in 8-character fragments.
bench_http.pyRound trip of POST /v1/evaluate over loopback, one keep-alive connection.

Run#

cd bench/
# engine time, call time, streaming and HTTP, three runs each
for s in bench_engine_time bench_evaluate_e2e bench_streaming bench_http; do
  for i in 1 2 3; do python $s.py; done
done
# cold start: five fresh processes
for i in 1 2 3 4 5; do python bench_cold_call.py; done

Each script prints the environment first, then one summary line per measurement:

cpu=Apple M5 Max cores=18 ram=48 GB os=macOS-26.6.2-arm64-arm-64bit python=3.11.4
evaluate() end-to-end wall clock: n=10000 p50=24.12us p90=26.75us p99=33.96us max=182.67us mean=24.00us ops/s=41,670
evaluate() engine-reported latency_us: n=10000 p50=18.50us p90=20.96us p99=26.00us max=144.08us mean=18.23us ops/s=54,840

macOS: numeric library conflict

If the import aborts with an OpenMP "duplicate library" message, run in a clean virtual environment, or set KMP_DUPLICATE_LIB_OK=TRUE for the benchmark run.

Compare and share#

  • Compare medians of three runs, not single runs. Expect p99 and maximum to vary most.
  • Measure your own payloads too: replace PAYLOADS in _common.py with a sample of real (test) traffic.
  • If your results differ a lot from ours on similar hardware, email hello@helixor.ai with the full script output.