# Decision Runtime benchmark results, 2026-09-27

Environment: Apple M5 Max (6 performance + 12 efficiency cores), 48 GB RAM, macOS 26.6.2,
Python 3.11.4, runtime 0.2.0 (pure-Python build), fastapi 0.123.0, uvicorn 0.40.0, pydantic 2.13.5.
Machine was NOT idle: load average 5.6-7.1 during all runs. Three runs per benchmark.

## examples/02_batch_benchmark.py (engine-reported latency_us; 5,000 calls)
| Run | ops/s | avg us | p50 | p90 | p99 |
|---|---|---|---|---|---|
| 1 | 40,243.9 | 18.11 | 17.04 | 21.75 | 30.62 |
| 2 | 42,490.4 | 17.44 | 17.00 | 19.71 | 33.38 |
| 3 | 44,309.4 | 16.80 | 16.92 | 19.00 | 21.50 |

## bench_evaluate_e2e.py (call time; 100 warm-up + 10,000 calls)
| Run | p50 us | p90 | p99 | max | mean | ops/s | engine p50 / p99 |
|---|---|---|---|---|---|---|---|
| 1 | 22.38 | 24.96 | 30.21 | 84.62 | 22.51 | 44,434 | 16.83 / 23.25 |
| 2 | 22.08 | 24.17 | 27.33 | 68.38 | 21.88 | 45,713 | 16.62 / 20.17 |
| 3 | 22.29 | 24.46 | 30.21 | 440.71 | 22.29 | 44,864 | 16.83 / 22.50 |

## bench_cold_call.py (5 fresh processes)
import 684-728 ms; construct 0.35-0.38 ms; first evaluate 354-366 us; second evaluate 38.6-41.0 us.

## bench_streaming.py (2,048-byte text, 8-char fragments, 20 streams)
| Run | push p50 us | p90 | p99 | max | 2 KB stream p50 |
|---|---|---|---|---|---|
| 1 | 18.79 | 21.08 | 26.00 | 96.42 | 4.96 ms |
| 2 | 19.21 | 21.62 | 25.92 | 68.29 | 5.03 ms |
| 3 | 19.17 | 21.50 | 25.54 | 97.75 | 5.06 ms |

## bench_http.py (POST /v1/evaluate, one keep-alive connection, 50 warm-up + 1,000 sequential)
| Run | p50 us | p90 | p99 | max | mean | req/s | server engine p50 |
|---|---|---|---|---|---|---|---|
| 1 | 336.25 | 359.54 | 386.25 | 494.33 | 337.75 | 2,961 | 28.12 |
| 2 | 340.42 | 363.29 | 436.50 | 659.92 | 342.30 | 2,921 | 27.88 |
| 3 | 339.33 | 360.17 | 377.62 | 454.33 | 340.04 | 2,941 | 28.54 |
