helixordevelopers

Sustainability

The most efficient decision is one that needs no accelerator, no network and no repeat. This page shows how to keep the runtime's footprint small, and how to measure it so you can prove it.

No model inference per decision#

A local evaluation runs a pack's extraction steps and rules on CPU, with no model inference. It uses no model tokens (tokens_spent is always 0) and no network transfer (egress_bytes is always 0). A decision that would otherwise need a language model call becomes CPU work measured in microseconds (latency_us, measured by bench_engine_time.py on the author's machine). No GPU or other accelerator is involved.

  • Decide locally first. Put the runtime in front of every model call it can answer or gate. For example, block a prompt that carries a card number before it is ever sent.
  • Escalate by rule. Send a request to hosted reasoning only when a local rule says it is needed; see Deployment patterns.
  • Keep GPUs for the work that needs them. Only the optional batch constraint solver can use a GPU. Per-request evaluation never does. Do not schedule decision services on accelerator nodes.

Right-size#

  • Measure one warmed-up worker for throughput and resident memory on the instance type you plan to use. Size from that, not from example values. See Cost optimization.
  • About one core per process. A single runtime process does not use more than one core. Allocating more to it wastes capacity.
  • Prefer the closest placement. In-process evaluation needs no extra container, no serialization and no network. A sidecar in every pod multiplies its memory by the pod count. A shared service pools capacity when many low-volume callers need decisions.
  • Use energy-efficient processors where your platform offers them. The runtime is pure Python and runs on x86_64 and arm64 (aarch64). Benchmark both on your workload and pick the one that does more work per watt.

Scale to zero where the pattern allows#

PatternIdle footprintHow to reduce it
In-process in a function or on-demand serviceNone when the platform scales the function to zero.Create the engine once per execution environment, not per request.
In-process in a long-running serviceShares the service's footprint.Nothing extra to do.
SidecarOne small container per pod, always on.Scale the pods with demand. Consider in-process for Python callers.
Shared serviceIts minimum replica count.Autoscale on CPU. Keep the minimum at what availability requires.
BatchNone between runs.Run batch jobs on demand or from an event, not on a timer.

A new process pays a one-time start cost: loading the pack and the first evaluation, which compiles patterns. Scale-to-zero trades that cost against idle capacity. Measure both before you choose.

Avoid redundant evaluation#

  • Evaluate once per payload and hop. If a gateway has already evaluated a payload and passed on its clean_text and receipt_hash, do not evaluate the same text again downstream unless a different pack applies.
  • Batch where you can. For records at rest, evaluate in a batch job instead of per request; see Batch.
  • Stream with care. A streaming session re-evaluates its held-back buffer on every push. Push token deltas as they arrive rather than re-sending the whole text, and use the default look-ahead unless you need a longer one; see Performance efficiency.
  • No polling. Check license expiry once a day and trigger everything else from events.
  • Log a fixed field set. One compact line per decision keeps storage and processing small; see Operational excellence.

Measure#

You cannot improve what you do not count. Track these alongside the metrics in Operational excellence:

MeasureFromUse it to
Decisions per CPU-secondDecision count divided by the CPU time of decision processesSpot regressions after pack or runtime changes.
Share of requests decided locallyLocal decisions divided by the sum of local and escalatedKeep model inference to the requests that need it.
Evaluations per requestEvaluation count divided by request countFind duplicate evaluation across hops.
Idle replica hoursReplica hours at low CPU utilizationTune minimum replicas and scale-to-zero.

For carbon reporting, use your platform's own emissions tooling. The runtime reports no energy or carbon figures.