Performance efficiency
Evaluation itself is fast. In most deployments, what you actually pay for is where the engine runs relative to the caller. Choose the closest placement that meets your other needs, and measure end to end on your own hardware.
Latency budget by placement#
| Placement | What you pay per decision | Use when |
|---|---|---|
| In-process library | Evaluation only. | The caller is Python and on the hot path. |
| Sidecar on loopback | Evaluation, plus request parsing, JSON serialization and a loopback round trip. | The caller is not Python, or you want the runtime in its own container. |
| Shared service in your network | All of the above, plus a network round trip, TLS and gateway authentication. | Many callers, central change control, and a latency budget in milliseconds. |
| Hosted reasoning API | An internet round trip plus multi-step reasoning. | Only for the hard minority of decisions; see Hosted escalation. |
Measure it yourself#
Run the engine-time benchmark (download it and _common.py from Reproduce). It evaluates 5,000 payloads in one process and reports throughput and P50, P90 and P99 latency:
python bench_engine_time.py
The results depend on your CPU and Python version. Figures quoted anywhere in these docs were measured with the scripts on Reproduce on the author's machine. Treat them as indications, not guarantees.
What latency_us covers
latency_us on each result measures evaluation inside the engine only. It excludes normalizing the input, building the result object, serialization, HTTP and the network. For a service, measure at the caller: wrap the client call in your own timer and export that as your latency metric.
To benchmark a deployment, repeat the measurement at each layer: in-process, then through the sidecar from the application container, then through the gateway. The differences tell you what each layer costs.
Warm up before serving#
The first evaluation in a process is noticeably slower than the rest, because it compiles the extraction patterns. Loading a pack also costs a license check and a decryption. Do both at start-up: load the engine and run one throw-away evaluation before you mark the process ready. See Reliability.
Input size and streaming#
- Evaluation cost grows with payload length. Every codon scans the whole text. Set an upper bound on payload size at your gateway, and split very large documents into records. See Batch.
- Streaming re-scans its buffer. Each chunk you push re-evaluates the whole held-back buffer, not just the new chunk. The buffer is released at delimiters once it is longer than the look-ahead window (28 characters by default), so cost per push stays small for ordinary text. Long runs without whitespace or punctuation keep up to twice the look-ahead window buffered and make each push more expensive. A larger
lookahead_charscatches longer values across chunk boundaries, but each push costs more. - Each push counts as one evaluation. This matters for decision counters and for Developer-tier rate limits.
- For real token streams over the network, use WebSocket.
POST /v1/evaluate/streamtakes a complete text and splits it on the server. It is useful for testing, but it cannot forward a live stream. The WebSocket endpoint acceptsstream_chunkmessages as tokens arrive. In Python, use a streaming session in-process. See Streaming.
Scale horizontally#
- Processes, not threads. Evaluation is CPU-bound Python. Run one worker process per core you want to use, each with its own engine.
- Replicas for the service.
helixor-pack serveis a single process. Scale it by adding containers behind your load balancer, sizing roughly one replica per core you allocate. - No shared state to coordinate. Instances share nothing at run time, so adding replicas needs no coordination. Keep your own aggregate counters in your metrics system.
- Memory. Each process holds the Python runtime, its dependencies and the unsealed pack. Measure resident memory of one warmed-up worker under load. Size requests and limits from that figure; the runtime does not publish one.
GPU acceleration#
Evaluation runs on CPU. The optional batch constraint solver can use a GPU through torch, and falls back to CPU when no GPU is present. It is used for large batch portfolio workloads, not for per-request evaluation, and it is not exposed by the decision service. Schedule GPU nodes only for those batch jobs.