Sustainability
The most efficient decision is one that needs no accelerator, no network and no repeat. This page shows how to keep the runtime's footprint small, and how to measure it so you can prove it.
No model inference per decision#
A local evaluation runs a pack's extraction steps and rules on CPU, with no model inference. It uses no model tokens (tokens_spent is always 0) and no network transfer (egress_bytes is always 0). A decision that would otherwise need a language model call becomes CPU work measured in microseconds (latency_us, measured by bench_engine_time.py on the author's machine). No GPU or other accelerator is involved.
- Decide locally first. Put the runtime in front of every model call it can answer or gate. For example, block a prompt that carries a card number before it is ever sent.
- Escalate by rule. Send a request to hosted reasoning only when a local rule says it is needed; see Deployment patterns.
- Keep GPUs for the work that needs them. Only the optional batch constraint solver can use a GPU. Per-request evaluation never does. Do not schedule decision services on accelerator nodes.
Right-size#
- Measure one warmed-up worker for throughput and resident memory on the instance type you plan to use. Size from that, not from example values. See Cost optimization.
- About one core per process. A single runtime process does not use more than one core. Allocating more to it wastes capacity.
- Prefer the closest placement. In-process evaluation needs no extra container, no serialization and no network. A sidecar in every pod multiplies its memory by the pod count. A shared service pools capacity when many low-volume callers need decisions.
- Use energy-efficient processors where your platform offers them. The runtime is pure Python and runs on x86_64 and arm64 (aarch64). Benchmark both on your workload and pick the one that does more work per watt.
Scale to zero where the pattern allows#
| Pattern | Idle footprint | How to reduce it |
|---|---|---|
| In-process in a function or on-demand service | None when the platform scales the function to zero. | Create the engine once per execution environment, not per request. |
| In-process in a long-running service | Shares the service's footprint. | Nothing extra to do. |
| Sidecar | One small container per pod, always on. | Scale the pods with demand. Consider in-process for Python callers. |
| Shared service | Its minimum replica count. | Autoscale on CPU. Keep the minimum at what availability requires. |
| Batch | None between runs. | Run batch jobs on demand or from an event, not on a timer. |
A new process pays a one-time start cost: loading the pack and the first evaluation, which compiles patterns. Scale-to-zero trades that cost against idle capacity. Measure both before you choose.
Avoid redundant evaluation#
- Evaluate once per payload and hop. If a gateway has already evaluated a payload and passed on its
clean_textandreceipt_hash, do not evaluate the same text again downstream unless a different pack applies. - Batch where you can. For records at rest, evaluate in a batch job instead of per request; see Batch.
- Stream with care. A streaming session re-evaluates its held-back buffer on every push. Push token deltas as they arrive rather than re-sending the whole text, and use the default look-ahead unless you need a longer one; see Performance efficiency.
- No polling. Check license expiry once a day and trigger everything else from events.
- Log a fixed field set. One compact line per decision keeps storage and processing small; see Operational excellence.
Measure#
You cannot improve what you do not count. Track these alongside the metrics in Operational excellence:
| Measure | From | Use it to |
|---|---|---|
| Decisions per CPU-second | Decision count divided by the CPU time of decision processes | Spot regressions after pack or runtime changes. |
| Share of requests decided locally | Local decisions divided by the sum of local and escalated | Keep model inference to the requests that need it. |
| Evaluations per request | Evaluation count divided by request count | Find duplicate evaluation across hops. |
| Idle replica hours | Replica hours at low CPU utilization | Tune minimum replicas and scale-to-zero. |
For carbon reporting, use your platform's own emissions tooling. The runtime reports no energy or carbon figures.