# Decision Runtime benchmark results, 2026-09-28

Environment: Apple M5 Max (6 performance + 12 efficiency cores), 48 GB RAM, macOS 26.6.2,
Python 3.11.4, runtime 0.2.1 (pure-Python build), fastapi 0.123.0, uvicorn 0.40.0, pydantic 2.13.5.
Machine was NOT idle: load average 7.5-10.5 during all runs (other test suites were running).
Three runs per benchmark, five fresh processes for cold start, run as described on the
Reproduce page. Previous results for runtime 0.2.0: results-2026-09-27.md.

## examples/02_batch_benchmark.py (engine-reported latency_us; 5,000 calls)
| Run | ops/s | avg us | p50 | p90 | p99 |
|---|---|---|---|---|---|
| 1 | 36,047.9 | 20.77 | 19.62 | 23.92 | 59.58 |
| 2 | 20,773.7 | 34.68 | 25.58 | 63.58 | 135.92 |
| 3 | 41,276.9 | 18.38 | 18.58 | 21.29 | 27.42 |

## bench_evaluate_e2e.py (call time; 100 warm-up + 10,000 calls)
| Run | p50 us | p90 | p99 | max | mean | ops/s | engine p50 / p99 |
|---|---|---|---|---|---|---|---|
| 1 | 24.12 | 26.75 | 33.96 | 182.67 | 24.00 | 41,670 | 18.50 / 26.00 |
| 2 | 24.88 | 32.29 | 48.83 | 6215.04 | 27.40 | 36,495 | 18.92 / 35.83 |
| 3 | 26.88 | 33.00 | 75.00 | 267.38 | 28.64 | 34,914 | 20.54 / 54.12 |

## bench_cold_call.py (5 fresh processes)
import 864-985 ms; construct 0.21-0.24 ms; first evaluate 107.5-112.6 us; second evaluate 48.0-55.1 us.

| Run | import ms | construct ms | first evaluate us | second evaluate us |
|---|---|---|---|---|
| 1 | 982.02 | 0.24 | 112.6 | 54.2 |
| 2 | 887.77 | 0.23 | 110.5 | 51.7 |
| 3 | 937.03 | 0.23 | 110.0 | 52.0 |
| 4 | 864.30 | 0.21 | 111.0 | 48.0 |
| 5 | 985.27 | 0.21 | 107.5 | 55.1 |

## bench_streaming.py (2,048-byte text, 8-char fragments, 20 streams)
| Run | push p50 us | p90 | p99 | max | 2 KB stream p50 |
|---|---|---|---|---|---|
| 1 | 57.92 | 75.75 | 159.42 | 307.62 | 15.11 ms |
| 2 | 59.46 | 76.96 | 159.67 | 280.04 | 15.89 ms |
| 3 | 53.83 | 66.46 | 130.88 | 175.75 | 14.06 ms |

## bench_http.py (POST /v1/evaluate, one keep-alive connection, 50 warm-up + 1,000 sequential)
| Run | p50 us | p90 | p99 | max | mean | req/s | server engine p50 |
|---|---|---|---|---|---|---|---|
| 1 | 340.17 | 369.08 | 487.67 | 1282.21 | 346.11 | 2,889 | 27.46 |
| 2 | 346.67 | 413.79 | 687.25 | 1559.08 | 365.05 | 2,739 | 27.96 |
| 3 | 359.62 | 416.04 | 648.12 | 1221.54 | 374.59 | 2,670 | 29.17 |

## Medians of three runs
| Measurement | p50 | p90 | p99 | Throughput |
|---|---|---|---|---|
| Engine time (latency_us) | 19.6 us | 23.9 us | 59.6 us | 36,048 evaluations/s |
| Call time, evaluate() | 24.9 us | 32.3 us | 48.8 us | 36,495 calls/s |
| Streaming, one push() of 8 characters | 57.9 us | 75.8 us | 159.4 us | 2 KB stream in 15.1 ms |
| HTTP round trip, loopback | 347 us | 414 us | 648 us | 2,739 requests/s |

Notes:
- Streaming push() is about three times the 0.2.0 figure. 0.2.1 keeps the buffer raw and
  evaluates candidate cuts before releasing text, so each push runs more than one evaluation.
- Import time still dominates cold start: the optional numeric library is still imported
  eagerly in 0.2.1 (lazy import did not land in this release).
- The first evaluate() in a process is about 110 us in 0.2.1 (354-366 us in 0.2.0).
- Engine run 2 and call-time runs 2-3 coincide with the highest load; compare medians, and
  expect tails to be noisier than on an idle host.
