Cost optimization
Evaluation uses no model tokens and needs no network. Your running cost is the compute you give the runtime, plus your license and any hosted reasoning you choose to use. This page shows how to keep each of those in proportion.
Where the cost goes#
| Cost | Driven by | How to control it |
|---|---|---|
| Compute | CPU time per evaluation, times decision volume, plus one runtime per process or replica. | Keep decisions in-process where you can. Size replicas from measurements. |
| Model tokens | None for local evaluation. tokens_spent is always 0. | Decide locally first. Escalate only what needs reasoning. |
| Data transfer | None for local evaluation. egress_bytes is always 0. | Keep the decision next to the data. Avoid cross-zone hops to a shared service. |
| License | Your tier and agreement. | Match the tier to what you use; see Licensing. |
| Hosted reasoning | Calls you escalate to the hosted API. | Escalate by rule, not by default; see below. |
| Operations | Logs, metrics, receipt storage and CI. | Log one compact line per decision. Store receipts, not payloads. |
Size compute from a measurement#
- Measure one worker
Run
bench_engine_time.pywith inputs that look like your traffic, on the instance type you will use. Note throughput per process and resident memory once warmed up. - Add the transport
If you use the decision service, repeat the measurement through it. Serialization and HTTP usually cost more than evaluation. See Performance efficiency.
- Compute replicas
Divide peak decisions per second by measured throughput per process. Add headroom for rolling updates and for the loss of one zone.
- Set requests from the measurement
Use the measured memory for container requests. Give each replica about one core; a single process does not use more.
Choose the cheapest placement that works#
- In-process adds no extra containers and no network hops. Use it for Python callers.
- Sidecar adds one small container per pod. It is cheap in latency, but its memory is paid in every pod.
- Shared service pools capacity across callers, which pays off when many low-volume callers need decisions. It adds gateway, network and TLS cost to every call.
Escalate deliberately#
Put the local runtime first on every path. Send a request to the hosted reasoning API only when a rule you wrote says the local decision is not enough, for example a specific action or an explicit abstain condition. Count escalations as a metric and review their rate. A rising escalation rate is a cost signal and usually a sign that a local rule is missing.
Keep operational cost flat#
- Log a fixed set of fields per decision; see Operational excellence. Payloads make logs large as well as unsafe.
- Sample debug-level output. Keep full receipts only for the retention you need.
- Use no scheduled polling. Check license expiry once a day and trigger everything else from events.