Governed execution
An agent that runs for a long time proposes its next action with a language model, so you cannot check that action in advance. Governed execution puts every such action through one action broker, under a policy the task declares, and sends any tool that runs code to an executor the policy chooses. This page explains the model, the policy, how an executor is chosen, what is refused, and what has been measured.
Status in one paragraph
The action broker and the policy-selected executors are available in the Helixor companion. The OS sandbox is Partial: it is Linux-only, so on macOS the default task is refused. SmolVM guest execution works on macOS arm64 for development only. The production SmolVM pool, an egress relay for sandboxed code, licence enforcement and the governing of provider CLI agents are Planned. See Status.
Why agent actions need governance#
A decision from the Decision Runtime is chosen from a legal set that the rules check before anything happens. A playbook move in a digital worker is chosen the same way. Neither needs isolation: the set of possible outcomes is known and checked in advance.
A long-running agent is different. A model proposes each next step: read this file, fetch this URL, run this command. You cannot list those steps ahead of time, and a command can do more than its name says. A package install can run a post-install script, and a test run can execute code the model wrote. So the control has to sit at the moment of action. It checks each action against a declared policy, records it, and contains any code the action runs.
Isolation here is a governance control. It is what lets you let an agent act without trusting each step it proposes.
The layered model#
There is one path. Every action goes through the broker. File and network actions run inside the broker, and no code runs for them. A tool that runs code goes to an executor, and the policy chooses the executor, never the caller.
| Layer | Threat | Contained by |
|---|---|---|
| Action | The model proposes a bad call: read ~/.ssh, write ~/.bashrc, post data to an unlisted host. | The broker: policy check and receipt for every action. |
| Code | An allowed action runs code that does something else, such as a post-install script. | The OS sandbox or a SmolVM guest. The broker cannot see inside a program it runs, so it never runs one without an executor. |
| Kernel | That code exploits a kernel bug through the system calls it is allowed to make. | A SmolVM guest only, because it has its own kernel. |
The action broker#
The agent never gets raw file or network access. It asks the broker to read_file, write_file, list_dir, http_fetch or run_tool. The broker checks each request against the task policy and runs it only if the policy allows it.
- Paths. Reads and writes are limited to workspace-relative prefixes. Absolute paths,
.., symlinks at any depth, hard links and denied names (.git,.env,.netrc,.sshby default) are refused. - Egress.
http_fetchreaches only hosts or CIDR ranges on the allow-list, on the listed ports and schemes. Every resolved address is checked. A host rule admits only globally routable addresses unlessallow_privateis set. The connection is pinned to the checked address, and redirects are not followed. - Limits. Read, write and response sizes, a fetch timeout, and actions per minute.
- Tools. Only declared tools run. A tool has a fixed program prefix, and the agent may add arguments only when the tool allows it.
The executors#
| Executor | What it is | Use it for |
|---|---|---|
| OS sandbox | A long-lived Linux process sandbox. The namespaces tier uses user, mount, pid, network and other namespaces, a fresh root with a read-only toolchain, cgroup v2 limits and seccomp. The landlock tier needs no privileges: Landlock file and network rules, seccomp, and per-process limits. Both run a per-session copy of the workspace with no network. The kernel is shared with the host. | A trusted toolchain for one tenant with no network: your own linters, formatters, test runners and build steps. |
| SmolVM guest | A micro-VM with its own kernel, one guest per session, destroyed on close. | Untrusted or customer-supplied code, multi-tenant pools where a kernel exploit would cross tenants, and code that needs network access by allow-list. |
The task policy#
Each task declares a policy. The contract is helixor.task_execution_policy.v1. A task policy overrides the worker's default policy, which overrides the built-in default. Unknown keys are refused.
The built-in default is the safe one: a trusted toolchain, one tenant, the executor chosen automatically, no broker egress, no network for code, and one declared shell tool.
| Field | What it declares |
|---|---|
workspace | Absolute path of the task workspace. Required. |
workload | trusted_toolchain (default) or untrusted_code. |
tenancy | single_tenant (default) or multi_tenant. |
executor | auto (default), os-sandbox, smolvm, or host (development profile only). |
os_sandbox_tier | auto, namespaces or landlock. An explicit tier is never downgraded. |
code_network | Network for code that tools run: off (default), restricted with allowed_cidrs, or open. Anything but off needs a SmolVM guest. |
egress | The broker's http_fetch allow-list. Each rule has a host or a cidr, ports, and optional schemes and allow_private. Empty means no egress. |
readable, write_back_paths, denied_names | Workspace-relative paths the broker may read, the paths written back when the session closes, and names always refused. |
tools | Declared tools: name, argv, allow_args, shell, timeout_seconds. |
limits | Memory, vCPUs, command timeout, session budget, workspace size, process count, read, write and fetch sizes, fetch timeout and actions per minute. |
secrets_allowlist | Names of secrets the task needs. A declared secret that is not supplied refuses the session. |
This policy runs a project's tests on a trusted toolchain. It lets the broker fetch from one API over HTTPS and write back only src and reports:
{
"schema_version": "helixor.task_execution_policy.v1",
"workspace": "/srv/agent/tasks/t-1042",
"workload": "trusted_toolchain",
"tenancy": "single_tenant",
"executor": "auto",
"code_network": { "mode": "off" },
"egress": [
{ "host": "api.example.com", "ports": [443], "schemes": ["https"] }
],
"readable": ["src", "tests", "docs"],
"write_back_paths": ["src", "reports"],
"tools": [
{ "name": "tests", "argv": ["python", "-m", "pytest", "-q"], "allow_args": true, "timeout_seconds": 300 }
],
"limits": {
"memory_mb": 1024,
"command_timeout_seconds": 300,
"session_budget_seconds": 1800,
"max_actions_per_minute": 120
},
"task_id": "t-1042"
}
This example validates against the schema. It was also passed to the companion's executor selection. On Linux, as an unprivileged user with no delegated cgroup, it selected the OS sandbox, landlock tier. On macOS it was refused, as described below.
How the executor is chosen#
With executor: auto, the workload, tenancy and code network decide:
| The policy declares | Executor | Tier |
|---|---|---|
trusted_toolchain, single_tenant, code network off | OS sandbox | namespaces when user namespaces and a writable delegated cgroup v2 subtree exist, otherwise landlock |
untrusted_code, or multi_tenant, or code network restricted or open | SmolVM guest | The VM backend |
executor: host | Host, no isolation | Development profile only; refused in production |
An explicit executor may add isolation. It may never remove isolation the workload needs. For example, executor: os-sandbox with workload: untrusted_code is refused with HELIXOR_EXECUTION_ISOLATION_POLICY_INVALID, because the task needs a VM.
Each session record and each receipt names the executor, its tier and the reason it was chosen.
Refusal behaviour#
If the executor the policy needs cannot start, the task is refused with HELIXOR_EXECUTOR_UNAVAILABLE and the facts that explain why. The task is never downgraded to a weaker executor. An unavailable OS sandbox is not swapped for a VM either: the policy says which executor it wants. The production profile has no host executor, so production never falls back to the host.
| Code | When |
|---|---|
HELIXOR_EXECUTOR_UNAVAILABLE | The selected executor cannot start on this host, or the production profile was asked for the host. |
HELIXOR_EXECUTION_ISOLATION_POLICY_INVALID | The policy is malformed, has unknown keys, or asks for less isolation than its workload needs. |
HELIXOR_EXECUTION_ISOLATION_CAPABILITY_UNSUPPORTED | The policy asks for something this host cannot enforce, such as a restricted code network where the CIDR allow-list cannot be enforced at the VM boundary. |
HELIXOR_EXECUTOR_NOT_ENTITLED | Reserved for the licence check before an executor starts. No licence is checked yet. |
HELIXOR_ACTION_NOT_DECLARED, HELIXOR_ACTION_PATH_DENIED, HELIXOR_ACTION_EGRESS_DENIED, HELIXOR_ACTION_LIMIT_EXCEEDED, HELIXOR_ACTION_RATE_LIMITED, HELIXOR_ACTION_POLICY_DENIED, HELIXOR_ACTION_SESSION_KILLED | The broker refused one action. The refusal is receipted like any other decision. |
helixor doctor lists each executor, whether it can run on this host and why, and what the default task policy does here. This is its execution section on a macOS arm64 development machine, run on 2026-09-29 against the merged companion code:
execution isolation (profile=development):
✓ broker in-process: available — file and network actions run in-process under the task policy; no code runs
✗ os-sandbox namespaces: unavailable — the OS sandbox is Linux-only (this host: darwin)
✗ os-sandbox landlock: unavailable — the OS sandbox is Linux-only (this host: darwin)
✓ smolvm qemu: available — smolvm 0.0.34 on darwin-arm64 (escape suite verified)
✓ host (no isolation): available — development profile only, and only when a task policy names executor 'host'; recorded as mode=host
! default task policy (broker + os-sandbox, no egress): code-running tools refused here — HELIXOR_EXECUTOR_UNAVAILABLE
the task needs the OS sandbox (trusted toolchain, single tenant, no code network) and no allowed tier is available here (namespaces: the OS sandbox is Linux-only (this host: darwin); landlock: the OS sandbox is Linux-only (this host: darwin)). Not downgrading to the host; declare executor 'smolvm' to use a VM.
- entitlement hooks: exec.brokered.v1=not_enforced, exec.sandboxed.v1=not_enforced, exec.isolated.v1=not_enforced
The default task is refused on macOS, which is the intended fail-closed result. A task that declares workload: untrusted_code or executor: smolvm runs in a SmolVM guest on this machine. The last line shows that every entitlement hook is a no-op that records enforced: false. Its three ids come from an earlier draft; the current draft replaces them with one entitlement with levels, described below. The companion also serves the same report at GET /api/v1/companion/execution_isolation/executors.
Receipts and the kill switch#
The broker appends a receipt (helixor.action_receipt.v1) for every decision, allowed or denied. The first receipt, session.open, records the executor selection and the entitlement decisions.
| Field | What it holds |
|---|---|
session_id, seq, action | Which session, the position in its chain, and the action. |
decision, reason_code | allowed, denied or failed, and why. |
policy_digest | SHA-256 of the policy the action was checked against. |
executor, executor_tier, executor_reason | The executor behind the session and why it was chosen. |
args_digest, result_digest | Digests of the arguments and the result. Raw payloads are never stored, so a receipt log is safe to ship to an audit store. |
at_unix_ms, duration_us | When the action ran and how long it took. |
prev_hash, hash | The SHA-256 of the previous receipt, and of this one. Removing, reordering or editing a receipt breaks the chain. |
What the chain does and does not prove
The chain is unkeyed SHA-256. It detects a changed, missing or reordered receipt inside a log. It does not stop someone who can rewrite the whole log from recomputing it. Keep the chain head (receipt_head on the session record) somewhere the agent cannot write. These action receipts belong to the companion. They are separate from the Decision Runtime's receipt_hash, which is unkeyed and not chained.
The kill switch ends a session at once. Every later action is refused with HELIXOR_ACTION_SESSION_KILLED, a running tool's executor is closed, and nothing is written back. A spent session budget triggers the kill switch too. In the Linux tests, a kill stopped a running sleep 300 in under 10 seconds and left no process behind.
In the companion, task sessions are opened with POST /api/v1/companion/execution_isolation/sessions and a task_policy body. The actions run, run_tool, read_file, write_file, list_dir, http_fetch, receipts, close and kill are posted to .../sessions/{id}/{action}. No endpoint runs code outside the broker.
Measured comparison#
These numbers come from Helixor's evaluation on 2026-09-29. It ran on a macOS arm64 development machine, with Linux runs inside a local container VM (aarch64, Landlock ABI 8). They are a development baseline. They were not measured on x86_64 or on production hosts.
| Broker | OS sandbox | SmolVM | |
|---|---|---|---|
| Escape suite | 11/11 broker actions; refuses unbrokered code | 10/10 | 10/10 |
| Start | 0 (in-process) | ~20 ms | ~1.3 s |
| Overhead | 8–31 µs per action | 2–3 ms per command | 2.5 ms per command |
| Memory | 0 | ~8 MB | ~400–720 MB |
| What defeats it | Code that a tool runs | A kernel exploit | A VM escape |
The escape suite tries ten breaches: reading host files, writing outside declared paths, planting symlinks through write-back, reading host credentials, reaching the network, seeing or signalling host processes, persisting after destroy, and exceeding memory, time and CPU limits. Run with no isolation, the host fails all ten, which proves the tests detect a breach. A plain container baseline passed 9 of 10.
The OS sandbox passed 10/10 in both tiers, including as an unprivileged user with no capabilities. The broker column is a separate suite of eleven attacks expressed as tool calls. The broker contains only actions that go through it. When the same calls were run through plain file and network APIs, they breached 8 of 8.
Platform notes#
- The broker is pure Python and runs anywhere the companion runs.
- The OS sandbox is Linux-only. The
namespacestier needs unprivileged user namespaces and a writable, delegated cgroup v2 subtree. In the evaluation, a default container security profile blocked it. Thelandlocktier needs Landlock ABI 6 or later and seccomp, and no privileges. It has no aggregate process or memory limit, so pair it with the container's own limits and one sandbox per container. The x86_64 system-call table has not been run yet; only aarch64 was measured. - SmolVM guests need hardware virtualization. Guest execution is enabled only on platforms where the escape suite has passed against a real guest. Today that is macOS arm64, with the SmolVM runtime version the suite passed with. Every other platform and runtime version refuses.
- The production SmolVM pool Planned needs Linux KVM hosts. A serverless container platform without nested virtualization cannot host it. It needs VMs with nested virtualization, Kubernetes node pools that allow it, or bare-metal hosts, and each needs its own escape-suite run before it is claimed.
- macOS process sandboxing is not a governed executor. It passed 8 of 10 escape tests, with no memory or CPU limit, so the policy never selects it.
Status#
| Part | Status | Scope |
|---|---|---|
| Action broker, receipts, kill switch | Available | In the Helixor companion, for every task session. |
| Policy-selected executors and refusal | Available | In the Helixor companion. Selection, refusal and the development-only host executor are covered by tests. |
| OS sandbox | Partial | Linux only. Shared kernel, trusted toolchain, no egress. On macOS the default task is refused. |
| SmolVM guest execution | Dev only | macOS arm64, for development. Not offered for production. |
| Production SmolVM pool on Linux KVM hosts | Planned | Nothing is provisioned. |
| Egress relay from sandboxed code through the broker | Planned | Sandboxed code has no network today. |
| Licence enforcement behind the entitlement hooks | Planned | The hooks run and record enforced: false. |
| Governing provider CLI agents | In progress | A provider CLI is itself the agent, launched as a host process, and its own tool calls do not go through the broker yet. The production profile refuses to launch one on the host. |
Entitlement levels#
Governed execution is licensed as one entitlement, exec.governed.v1, with a level. The levels are cumulative. This is a draft under review: it is not in any licence yet, and nothing checks it.
| Level | Adds | Issuable | Evidence required |
|---|---|---|---|
brokered | Every file and network action through the broker, with receipts, limits and the kill switch. Contains brokered actions only, never code a tool runs. | Preview | The broker-action suite. |
sandboxed | Code-running tools in the OS sandbox. | Partial Scope: shared kernel, trusted toolchain, no egress. | The escape suite on the target platform. |
vm | Code-running tools in a SmolVM guest from the production pool. | Not available Until the production pool exists. | The SmolVM escape suite run on the production pool. |
Each level is issuable only on its own evidence: its escape suite, run on the target platform. A run on a developer laptop does not count. A licence never names more isolation than was built and measured, so a shared-kernel sandbox is never sold as VM isolation. See the capability and licensing table.