Deployment patterns
Five reference topologies cover most production deployments. Start with the first one that fits: each later pattern adds a hop, a component or a boundary to operate.
Reference examples, not shipped artifacts
The Dockerfile, pod spec and network policy on this page are examples for you to adapt. Helixor does not publish a container image or deployment manifests for the runtime. Paths, ports, labels and names are illustrative.
| Pattern | Callers | Extra hop | You operate |
|---|---|---|---|
| 1. In-process library | Python | None | Nothing extra |
| 2. Sidecar on loopback | Any language, same pod | Loopback HTTP | One container per pod |
| 3. Shared in-network service | Many services | Network + gateway | A service, a gateway, network policy |
| 4. Air-gapped | Any | As 1–3 | Offline delivery of wheels, licenses and packs |
| 5. Hybrid escalation | Any | Internet, for escalated requests only | Escalation rules and data-release policy |
1. In-process library#
Use when the caller is written in Python and the decision is on its hot path.
Trade-offs. This is the lowest latency and the least infrastructure. A rule change means restarting your application. The runtime shares the application's privileges and its network access.
from helixor_runtime import HelixorEngine
engine = HelixorEngine.load_pack("/run/helixor/pack/guard.hxpack",
license_file="/run/secrets/helixor/helixor.lic")
engine.evaluate("warm-up") # compile patterns before serving
def handle(text: str) -> str:
try:
result = engine.evaluate(text)
except Exception:
raise PermissionError("blocked: decision unavailable") # fail closed
if result.action.startswith("block_") or any(t.severity == "FATAL" for t in result.triggers):
raise PermissionError(result.reason)
return result.remedy.clean_text
Create the engine once per worker process, at start-up, not per request. With a pre-fork server, create it after the fork.
HelixorEngine() with no arguments runs the built-in example pack instead. Check engine.pack_id at start-up; see Reliability. (In 0.2.0, load_pack could not load sealed packs, so this pattern was limited to the built-in pack; 0.2.1 fixes that.)
2. Sidecar decision service on loopback#
Use when the caller is not Python, or you want to ship rule changes without rebuilding the application image.
Trade-offs. You pay a loopback round trip and JSON serialization per decision, and one sidecar's memory in every pod. The loopback bind is your main access control. Add a service token (HELIXOR_SERVICE_TOKEN) if other processes share the pod's network namespace.
Hardened sidecar image#
FROM python:3.11-slim
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1
# Unprivileged user; empty working directory so no stray .env files are read
RUN useradd --system --uid 10001 --create-home --home-dir /home/helixor helixor
# The wheel that came with your evaluation access, pinned by checksum in your build
COPY helixor_runtime-0.2.1-py3-none-any.whl /tmp/
RUN pip install --no-cache-dir /tmp/helixor_runtime-0.2.1-py3-none-any.whl \
&& rm -f /tmp/*.whl
# Evaluation probe from the Reliability page
COPY probe.py /opt/helixor/probe.py
WORKDIR /home/helixor
USER 10001
EXPOSE 18734
# Pack and license are mounted read-only at run time, never copied into the image
ENTRYPOINT ["helixor-pack", "serve", \
"--pack", "/run/helixor/pack/guard.hxpack", \
"--license", "/run/secrets/helixor/helixor.lic", \
"--host", "127.0.0.1", "--port", "18734"]
HEALTHCHECK --interval=30s --timeout=3s --start-period=10s --retries=3 \
CMD ["python", "/opt/helixor/probe.py"]
probe.py is the evaluation probe in Reliability. It checks the pack ID and a real verdict, which GET /v1/health does not.
Pod spec sketch#
apiVersion: apps/v1
kind: Deployment
metadata:
name: claims-api
spec:
replicas: 3
selector:
matchLabels: { app: claims-api }
template:
metadata:
labels: { app: claims-api }
spec:
securityContext:
runAsNonRoot: true
fsGroup: 10001
containers:
- name: app
image: registry.example.com/claims-api:2.4.0
env:
- name: DECISION_URL
value: "http://127.0.0.1:18734"
- name: helixor-decision
image: registry.example.com/helixor-decision:0.2.1
securityContext:
runAsUser: 10001
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
capabilities: { drop: ["ALL"] }
resources:
requests: { cpu: "1", memory: "256Mi" } # replace with your measurement
limits: { cpu: "1", memory: "512Mi" }
volumeMounts:
- { name: pack, mountPath: /run/helixor/pack, readOnly: true }
- { name: license, mountPath: /run/secrets/helixor, readOnly: true }
- { name: tmp, mountPath: /tmp }
startupProbe:
exec: { command: ["python", "/opt/helixor/probe.py"] }
periodSeconds: 2
failureThreshold: 30
readinessProbe:
exec: { command: ["python", "/opt/helixor/probe.py"] }
periodSeconds: 10
livenessProbe:
exec: { command: ["python", "/opt/helixor/probe.py"] }
periodSeconds: 30
volumes:
- name: license
secret:
secretName: helixor-license-2026 # versioned: license + packs roll together
defaultMode: 0440
items: [{ key: helixor.lic, path: helixor.lic }]
- name: pack
configMap:
name: claims-guard-pack-1-4-0 # binaryData; or an init container that
items: [{ key: guard.hxpack, path: guard.hxpack }] # fetches from your artifact store
- name: tmp
emptyDir: {}
- Give the secret and the pack versioned names. A rollout then swaps both together, and a rollback swaps both back.
- The memory values are placeholders. Size them from a measured, warmed-up worker; see Cost optimization.
- A network policy applies to the whole pod, not to a single container. In this pattern, egress rules must allow what your application needs. For a decision workload with egress fully denied, use pattern 3.
3. Shared in-network decision service#
Use when many services need the same decisions, you want one place to roll out rule changes, or a central team owns the policies.
Trade-offs. This pattern adds a network round trip, TLS and a gateway to every decision, and it is a shared dependency to keep available. In exchange, you get pooled capacity, one rollout per rule change and one audit point.
- Keep
helixor-pack servebound to127.0.0.1. Put an authenticating proxy in the same pod: a mesh sidecar with mutual TLS and an authorization policy, or a reverse proxy that checks a token and terminates TLS. Only the proxy port is reachable. SettingHELIXOR_SERVICE_TOKENon the service as well gives a second check behind the proxy. - At the gateway, enforce a maximum body size, a timeout and CORS for the exact origins you serve.
- Callers treat a timeout or a 5xx from the gateway as a block.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: helixor-decision
namespace: decisions
spec:
podSelector:
matchLabels: { app: helixor-decision }
policyTypes: ["Ingress", "Egress"]
ingress:
- from:
- namespaceSelector:
matchLabels: { role: gateway }
ports:
- { protocol: TCP, port: 8443 } # the authenticating proxy, not 18734
egress: [] # deny all egress, including DNS
If your mesh sidecar needs to reach its control plane, add that single destination to egress. Nothing else is required: the runtime makes no outbound calls once the pack and license are mounted.
4. Air-gapped#
Use when the site has no route to the internet.
Trade-offs. Every renewal and every catalog pack update needs a manual transfer. Remote compilation, the catalog and escalation are unavailable.
- Bring in the runtime
Transfer the wheel and its checksum. Serve it from your internal package mirror, and build the images from patterns 1–3 inside the site.
- Bring in the license
Transfer the
.hxlicand load it into the site's secret store. To compile your own playbooks inside the site, the license must include the custom playbook compilation feature. - Compile locally
Run
helixor-pack compilewithout--remote. Download catalog packs withdownload-packoutside the site, for the same license, and transfer them in. - Block the outbound commands
Inside the site, never use
compile --remote,catalogordownload-pack. They contact the Helixor compiler service and fail without a route. - Keep time
License validity uses the local clock. Sync every host to the site's internal time source.
- Plan renewals
Transfer the renewed license before the end date. Then rebuild every pack for it inside the site, as in Reliability.
5. Hybrid escalation to hosted reasoning#
Use when a minority of decisions need evidence you do not hold locally, multi-step reasoning, or a calibrated probability of being correct.
Trade-offs. Escalated requests cost an internet round trip and hosted usage. They also send data outside your perimeter, so they need a data-release policy. The local decision still runs first, and a block is final.
- Escalate by rule. For example, escalate only when the local action is a permit and the request type is on an allow-list. Never escalate a payload the local runtime blocked.
- Send only redacted text. Use
result.remedy.clean_text, never the original payload. - Authenticate. Send your API key in the
X-Helixor-API-Keyheader, from your secret manager. - Fail closed. If the hosted call times out or fails, treat the request as undecided and apply your fallback policy, which should be a block or a human review.
- Log the link. Record the local
receipt_hashnext to the hosted response's proof identifier, so an auditor can follow one request across both.
import os, httpx
API_URL = os.environ["HELIXOR_API_URL"]
API_KEY = open("/run/secrets/helixor/api-key").read().strip()
def decide(engine, text, request_type):
local = engine.evaluate(text)
if local.action.startswith("block_") or any(t.severity == "FATAL" for t in local.triggers):
return {"final": "block", "receipt": local.receipt_hash}
if request_type not in ESCALATE_TYPES: # your allow-list
return {"final": "allow", "text": local.remedy.clean_text, "receipt": local.receipt_hash}
body = build_decide_request(local.remedy.clean_text) # goal, questions, evidence
try:
r = httpx.post(f"{API_URL}/v1/decide", json=body,
headers={"X-Helixor-API-Key": API_KEY}, timeout=5.0)
r.raise_for_status()
except httpx.HTTPError:
return {"final": "review", "receipt": local.receipt_hash} # fail closed
return {"final": "hosted", "hosted": r.json(), "receipt": local.receipt_hash}
A /v1/decide request needs a goal and a list of questions, with the evidence you choose to share. Hosted escalation describes the request and response. An escalation contract built into the runtime is Planned. It would make the local result carry a typed escalation request and link the two proofs. Until it ships, the logic above lives in your code.