How agents work
An agent pursues a goal by repeating one loop: perceive what is happening, choose a strategy the rules allow, act, and learn from how things turned out. This page walks the loop step by step and shows, for each step, the mechanism the Helixor sales assistant uses today and whether it is shipped or planned.
The loop#
The loop is the same for any agent. What makes a Helixor worker different is where each step's decision is made. Steps that decide are declared rules or small trained heads that can abstain. A language model is used only to read free text into proposals and to word replies, and its output is checked before it is used.
Goals#
The idea. An agent needs to know what counts as done, and how to tell whether it is getting closer. Without that, it can only answer questions.
In the sales assistant Preview. The goal is to qualify a visitor to one next step. The outcome set is fixed: start a trial, book a demo, talk to an engineer, come back later, or stop. Progress is the qualification record: as slots fill, the conversation moves from discover to qualify to propose. A rule table picks the outcome from the record and gives a reason. When the record is too thin, the decision names the slots that block it, and the worker asks about those next.
Perception#
The idea. The agent turns what it observes into something it can decide on, and knows how sure it is.
In the sales assistant Preview. Each message goes through three readers, in this order:
- Scope check. Deterministic, no model. The message is checked against the declared product and playbook packs. An out-of-scope request is declined from a rule template before any model is called. A request that needs an input the worker does not have yet (a demo without a known use case, for example) turns the move into asking for it.
- Intent and sentiment heads. Small trained classifiers, not a language model. Each returns a choice or a score with a confidence, and abstains below its threshold. An unsure read on an intent that would end the conversation is not acted on. The heads were trained on synthetic conversations, and their calibration has not yet been validated against real outcomes.
- Verbatim spans. A language model proposes what the visitor said about themselves. A value is stored only if it appears word for word in the message; anything inferred is rejected, and the rejection is recorded.
With the visitor's consent, what the site already knows (for example the landing page) is used as a hypothesis that shapes the first question. It never fills a slot.
Strategy selection under constraints#
The idea. The agent has several techniques. Some are never acceptable in a given situation, whatever they might achieve. Constraints decide what is legal; a preference chooses among what is left.
In the sales assistant Preview. Nine moves exist, from answering plainly to asking for a signup. The playbook turns intent and sentiment into the legal set, with hard rules first: hand off when asked, stop after a no, no selling to a hostile visitor, no signup ask at low sentiment or too soon after the last one. Once the visitor is qualified, the closing move for the decided outcome becomes legal and further discovery questions are removed. A trained move head then scores only the legal moves. One legal move is forced. When the head is unsure, the playbook's default is used and the turn is marked escalate. The legal set and the rule that shaped it go into the turn record.
Action#
The idea. The agent does something in the world, and what it does must be checked before it lands.
In the sales assistant Preview. The action is a reply and, sometimes, a next step:
- A language model writes the reply for the chosen move, using only facts retrieved from the public knowledge base, listed to it as short keys.
- The reply gate checks the draft: every citation resolves to a fact retrieved this turn, no price or guarantee, no excluded name, one question at most. A refused draft goes back once with the reasons; a second refusal sends a hand-off line and queues the question for a person.
- Hand-offs and stops come from rule templates and end the conversation.
- For a demo or engineer next step, the worker offers a booking card. Booking is disabled in production today, so it offers the contact path instead of times.
Every turn is then written to the conversation record: the reads and their confidence, the scope verdict, the legal set and rule, the chosen move and why, the cited facts, the gate result and the version of each head.
Outcome learning#
The idea. The agent gets better because it compares what it decided with what happened, and changes only through a gate that proves the change helps.
In the sales assistant Planned. The worker does not learn from outcomes yet. What exists today is the material learning would use:
- A conversion forecast is recorded on every turn. It comes from a head trained on synthetic conversations and has not been validated.
- Questions the worker could not answer, and requests the packs did not cover, are recorded on each conversation.
- Every turn names the head versions that produced it, so a later outcome can be traced to the exact heads and rules that were running.
The planned path: labels from reviewed real conversations, calibration measured against what actually happened, and a retrained head run in shadow beside the pinned one, promoted as a new version only when an evaluation gate passes. The Decision Runtime's outcome memory is available on its own; the sales assistant does not use it.
Shipped and planned, at a glance#
| Step | Sales assistant mechanism | Language model | Status |
|---|---|---|---|
| Goal | Fixed outcome set, qualification record, stage and outcome rules | No | Preview |
| Perception | Scope check; intent and sentiment heads with abstention; verbatim spans | Proposes spans only | Preview |
| Strategy | Playbook legal set and hard rules; move head over the legal set | No | Preview |
| Action | Worded reply from retrieved facts; reply gate with one repair; templates for hand-offs | Writes the draft | Preview |
| Record | Durable conversation record, one row per turn | No | Preview |
| Outcome learning | Forecasts and gaps recorded; no retraining from outcomes | No | Planned |
Where governed data and isolated execution fit#
Two further capabilities extend what a worker can perceive and do. The sales assistant uses neither. It perceives only the visitor's messages and its public knowledge base, and it runs no commands.
- Governed data Partial would widen perception: a worker could read live rows from your systems under ontology row and field policy, with no copy made. Today that covers virtual reads with filter and projection pushdown, no aggregation, and joins capped at 2,000 rows.
- Isolated execution Not available would widen action: a worker could run commands or code in an isolated guest. It is not built, and nothing runs in isolation today. It ships only with a test proving a command in the guest cannot touch the host.
Scope, status and draft licensing for both are in the capability table.
Related#
Build a digital worker
What a worker is, how it differs from the embedded runtime, and how it is licensed.
Tutorial · PreviewHow the sales assistant was built
Every mechanism on this page, with the code behind it.
Try itTalk to the sales assistant
Watch perception, strategy and action turn by turn.