How the sales assistant was built
This walkthrough follows the build of the Helixor website sales assistant, in the order the work happened. It shows how the platform's parts fit into one conversational agent: decision heads read intent and pick moves, a playbook bounds which moves are legal, a knowledge base supplies the only facts the agent may state, and gates check every reply. Sales is the worked example. The same pattern applies to any agent that has to reach an outcome through a conversation.
Status
The sales assistant runs as Preview. It has not launched publicly. Everything below marked Preview exists and is tested. Items marked Planned are designed but not built. The code excerpts are shortened from the worker's source, and the names are the real ones.
What you'll learn#
- How to structure an agent so that every choice is a decision you can inspect, and the language model only writes the words.
- How decision heads read intent and sentiment, and why some reads must be confident before the agent acts on them.
- How a qualification record and an outcome decision make the agent pursue a goal rather than only answer questions.
- How a playbook limits which techniques are legal on each turn, and how a head chooses among them.
- How grounding, a reply gate and a single repair pass keep replies true without handing everything to a person.
- How to engineer against a prototype with a turn-by-turn parity oracle, and what evaluation has to show before a head is trusted.
1. The goal: every choice is a decision#
The brief was a conversational digital employee. It answers a website visitor's questions, finds out who they are and what they need, and drives toward one concrete next step: start a trial, book a demo, talk to an engineer, come back later, or stop.
The design rule came before any code. Every choice the agent makes is a decision, and the language model is only the voice. Four kinds of choice happen on every turn, and none of them is left to the model:
| Choice | Made by | What the model does |
|---|---|---|
| What the visitor means, and how they feel | Decision heads for intent and sentiment | Nothing |
| What we have learned about them | A substring gate over spans the model proposes | Proposes spans |
| Which technique to use now, and which next step to aim for | Playbook rules, then the next-move head over the legal moves | Nothing |
| Whether a reply may be sent | The reply gate | Writes the draft, and one repair if the gate refuses it |
This gives every turn a record of why it happened. It also means a new head, a new rule or a new fact changes behavior in a way you can test, not a prompt you have to re-tune.
2. The agent development model#
The assistant is a conversational employee: a digital employee kind whose runtime is a turn loop, not a batch decision or a producer job. Like every employee, it is declared in a manifest. The manifest pins the four heads it reads with, names its knowledge base and playbook, and sets the outcomes it may reach.
schema: helixor.employee_manifest.v1
employee_id: sales_rep
kind: conversational
conversational:
playbook_id: sales_rep.v1
goal_id: sales.qualify_to_next_step
kb_refs: [kb.sales.collateral]
actor_clearance: public
retrieval_limit: 10
# Pinned admitted heads. A retrained head is a new pin.
heads:
intent:
family_id: sales.prospect_intent
digest: sdigest:8203dd30278aaf5cea90
sentiment:
family_id: sales.prospect_sentiment
digest: sdigest:3c5dbbdda0bf2c1498bf
move:
family_id: sales.next_move
digest: sdigest:2eac9f2efc6cec3e7b27
convert:
family_id: sales.will_convert
digest: sdigest:ae7778580fd5d9816e40
outcomes: [start_trial, book_demo, hand_to_engineer, nurture, disqualify]
autonomy_level: act_within_playbook
signup_ask_cooldown_turns: 3
Manifest validation fails closed. A conversational employee with no heads is refused, and so is one that also declares a batch decision runtime. Loading each head is fail-closed too. The artifact must be admitted, its family must match the pin, its digest must match the pin, and where the artifact carries enough to recompute the digest, the recomputed value must match as well:
if artifact.get("admission") != "admitted":
raise HeadLoadError("HEAD_NOT_ADMITTED", ...)
if artifact.get("family_id") != ref.family_id:
raise HeadLoadError("HEAD_FAMILY_MISMATCH", ...)
if artifact.get("digest") != ref.digest:
raise HeadLoadError("HEAD_DIGEST_MISMATCH", ...)
recomputed = artifact_digest(artifact)
if recomputed is not None and recomputed != ref.digest:
raise HeadLoadError("HEAD_DIGEST_FORGED", ...)
Retraining a head never changes the running agent silently. A new head is a new digest, and a new digest is a manifest change that goes through review.
The runtime is one method. handle(state, text) takes the conversation state and the visitor's message and returns a TurnResult: the decision with all of its inputs, the reply, the cited fact ids, where the words came from (model, rule template or gate), and any next-step surface to show.
def handle(self, state: ConversationState, text: str) -> TurnResult:
read = self.read_turn(text) # heads
learned = self.learn_slots(state, text) if read["level"] >= 1 else {...}
turn = self.decide_turn(state, text, read) # playbook + move head
...
out = self.generate_reply(state, turn) # model words + gate
...
return TurnResult(decision=turn, reply=out["reply"], fact_ids=..., source=out["source"],
gated=..., step=self.step_for(state, turn), cta=self.cta_for(turn), ended=state.ended)
Each conversation is a durable record (helixor.sales_conversation.v1) with the visitor context, the worker's state and one row per turn. Each row stores the intent and its confidence, whether the head abstained, the sentiment level, the stage, the legal moves and the rule that bounded them, the chosen move and why, the outcome decision, the retrieved and cited facts, the gate result, and the digest of every head that took part. The store checks the revision before writing and refuses a stale write with CONVERSATION_CONFLICT. The website widget reaches the service through its own server route, so the browser never holds the service credential.
3. Intention: reading the visitor with decision heads#
Four decision heads do the reading. Each is a small admitted model with a fixed option set, a calibrated distribution and a derived abstention threshold. They are the same kind of head used elsewhere on the platform. Here they score hashed character n-grams with a softmax, and confidence is one minus the normalized entropy.
| Head family | Reads | Options |
|---|---|---|
sales.prospect_intent | The visitor's message | pricing_question, feature_question, price_objection, timing_objection, trust_objection, ready_to_sign, not_interested, wants_human, small_talk, situation_answer |
sales.prospect_sentiment | The visitor's message | An ordinal 0 (hostile) to 4. The level is the rounded expected value. |
sales.next_move | The conversation prefix, tagged with stage, intent and sentiment | Nine techniques, scored only over the moves the playbook allows this turn |
sales.will_convert | The same prefix | true or false: a forecast recorded on every turn |
The will-convert head is trained with prefix labelling. Every prefix of a conversation is labelled with that conversation's final outcome. A forecast made at turn two is judged by how the conversation ended, not by a guess about turn two.
situation_answer was added during the build. The first heads read a visitor answering a discovery question ("we're a platform team of about forty") as an objection, because nothing in their option set described an answer. An agent that asks questions needs an intent for the answers.
Abstention, and why terminal intents need a confident read#
A head abstains when its confidence is below its threshold. On most intents, an abstained read still picks the playbook row, and the turn record says that the head was unsure. Two intents are different. not_interested and wants_human end the conversation by hard rule, so the worker acts on them only when the read is confident:
if not intent.unavailable and intent.abstained and intent_id in pb.TERMINAL_INTENTS:
# a terminal intent (its row ends the conversation by hard rule) needs a confident read
nxt = next((o for o, _ in intent.ranked if o not in pb.TERMINAL_INTENTS), None)
guarded, intent_id = intent_id, (nxt or "small_talk")
This came from a real failure. A visitor asked how Helixor differs from another tool. The intent head was unsure, its top option happened to be not_interested, and the conversation ended politely on a buying question. The general lesson: the cost of acting on an unsure read depends on what the action does. An irreversible action needs a stronger read than a reversible one.
4. Goal-oriented decision making#
The first prototype only answered questions. That is not selling. The agent needed a goal and a way to measure progress toward it. Three pieces provide that.
The qualification record#
The record has eight slots, asked in this order: role, company, use_case, pain, scale, current_approach, timeline, constraints. Who the visitor is comes first, then what brought them, then what it costs them and how big it is.
Filling a slot is propose-and-admit. The model reads the message and proposes values. A value is stored only if it appears word for word in the visitor's message. Anything paraphrased, inferred or phrased as a question is rejected, and the turn records what was rejected.
def verify_slots(proposed, text):
"""Propose-only. A value is admitted only as a verbatim span of the message."""
low = str(text or "").lower()
for k in SLOT_KEYS:
v = ... # the proposed value, stripped
if len(v) <= 160 and v.lower() in low and not v.endswith("?"):
admitted[k] = v
else:
rejected.append(k)
The record therefore holds only what the visitor said. Nothing the model inferred about them is ever stored as a fact.
Stages and the outcome decision#
The stage follows from the record. A conversation is in discover until the use case is known, then qualify. It moves to propose when the use case and two of scale, pain and timeline are known and sentiment is at least neutral. Objection intents move it to objection, and a ready-to-sign intent moves it to close.
The outcome is a choice over a fixed set: start_trial, book_demo, hand_to_engineer, nurture, disqualify. It is made by a rule table over the record, and each outcome carries its reason:
- A stated deployment or compliance constraint (air-gapped, on-premises, regulated, data residency) goes to
hand_to_engineer. - A far-off timeline goes to
nurture. - A team rollout (200 or more people, or company-wide) goes to
book_demo. A demo on their own material comes before a signup. - A self-serve product at a smaller scale goes to
start_trial. Another product goes tobook_demo. - A no goes to
disqualify.
When the record is too thin, the decision says so and names what is blocking it (for example "use_case" or "scale or timeline"). That blocker is what the agent asks about next.
The per-turn drive#
Every conversational turn ends with a drive. It is either one question for the next unfilled slot or, once an outcome is decided in the propose or close stage, the decided next step. The worker never drives when sentiment is at 1 or below. If the model's reply leaves out the drive, the worker appends it from a rule template, and the turn's source records it, for example question about scale added by rule template.
A known visitor context can turn an open question into a confirmation. Someone who arrived on a product page is asked whether that product is what they came for, instead of an open "what brought you here".
5. Strategy selection: a playbook bounds the moves#
The agent has nine techniques: answer_plainly, discovery_question, social_proof, reframe_value, handle_objection, offer_trial, ask_for_signup, hand_to_human and disengage. Choosing among them happens in two steps: the playbook decides what is legal, then the next-move head chooses among the legal moves.
The playbook is a matrix from intent to allowed moves, plus hard rules applied over it:
| Rule | Effect | Status |
|---|---|---|
| Disclose that it is an assistant | The opening line says so, and the first model-worded reply is told to say so. | Preview |
| Hand off on request | A wants_human intent allows only hand_to_human. The conversation ends with the contact path. | Preview |
| Stop after a no | A not_interested intent allows only disengage. | Preview |
| No selling to a hostile visitor | At sentiment 0, only answering, handing off or disengaging is allowed. | Preview |
| No signup ask at low sentiment, and not too often | At sentiment 1 or below, or within three turns of the last ask, ask_for_signup is removed. At sentiment 1, social proof is removed and a hand-off is always offered. | Preview |
| No prices, discounts, guarantees or trial terms | The reply gate refuses them. No public source carries them, and pricing goes to a person. | Preview |
| No payment details in the conversation | The booking step collects only a name and a work email. | Preview |
| Discounts through a human approval step | A discount request routed to an existing approval step instead of being refused. | Planned |
| Scarcity only with a grounded offer | An urgency technique allowed only when a real, cited offer exists. | Planned |
| Playbook as an admitted rule pack | The same rules compiled and admitted as a versioned pack, instead of the worker's declared tables. | Planned |
Each call returns the legal set and the rule that shaped it, and both go into the turn record:
def legal_moves(intent, sentiment_level, signup_asked_recently):
if intent == "wants_human":
return ["hand_to_human"], "hard rule: the prospect asked for a person"
if intent == "not_interested":
return ["disengage"], "hard rule: after a no, only disengage"
if sentiment_level == 0:
return ["answer_plainly", "hand_to_human", "disengage"], "hard rule: hostile sentiment — no selling moves"
moves = list(_BASE_MATRIX.get(intent, _DEFAULT_ROW))
...
The stage and outcome then adjust the set. Once qualified with a decided outcome, the closing move for that outcome becomes legal and further discovery questions are removed. The next-move head scores only the legal moves. If one move is legal, it is forced. If the head is unsure, the turn uses the playbook's default move and the verdict is escalate. A head can prefer a technique, but it can never pick one the playbook forbids.
6. Grounding: a knowledge base of public collateral#
The agent may state only what the knowledge base says. The sales knowledge base is built from public collateral only: site pages, published whitepapers and user documentation. Internal material and customer presentations are excluded at build time. Every document is ingested strictly and losslessly, so every fact keeps a citation id, its product, its source and, for measurements, its caption.
Retrieval is lexical: idf-weighted token overlap with a boost for the product the message names. Nothing is embedded and nothing is learned. When a short or generic message matches nothing, the worker offers the overview facts instead and records fallback: overview. It does not answer from nothing.
Short keys for the model, citation ids for the record#
An early version refused every cited reply. The model was given long citation ids and returned only part of each one, so an exact-match check failed every time. The fix is general: give the model short keys and map them back. The prompt lists facts as F1 to Fn. The gate resolves a key, a full id or a unique suffix to the citation id retrieved this turn. Anything else is an unknown citation, and unknown citations fail the gate.
F1 [<product>; <source>]: <fact text> F2 [<product>; <source>; measurement: <caption>]: <fact text>
Claims exclusions and absolute answers#
The knowledge base never stores the list of names that must not appear. It records which claims policy pack it was built against, with that pack's version and digest. At load time the index resolves the pack again and refuses a digest mismatch. The reply gate refuses any reply that contains a name on the list. When a visitor asks how Helixor differs from something else, the agent describes what Helixor does in absolute terms from the facts and does not repeat the other name.
7. The reply gate, and one repair#
A drafted reply is sent only if it passes every check:
- It has text, and every cited key resolves to a fact retrieved this turn.
- It cites at least one fact. Small talk answered plainly is the only exception.
- It names no excluded company or product.
- It contains no price, discount, guarantee, trial-term, per-seat, per-user or percentage wording.
- It asks for a signup only when the move is
ask_for_signupor the decided next step is a trial. - It asks at most one question when the drive is a question, and it stays within the length limit.
Each failed check produces a reason, such as no facts cited or asked more than one question.
Fixing "punts too often"#
The first widget build handed the visitor to a person far too often. The causes, found by reading turn records, were all general:
- Two instructions contradicted each other. The discovery move said "ask one question", and the drive added another question. Almost every such reply failed the one-question check. Now a discovery move that carries a drive question is told that the drive's question is the only one.
- A refused draft was final. Now the gate's reasons go back to the model once, with the refused draft, and the model rewrites it under the same move. A second refusal sends the escalation line and puts the question in the unanswered queue for a person.
- Lexical misses on short messages. These now fall back to the overview facts, as described above.
- Unsure terminal reads ended conversations. These now require a confident read, as described in section 3.
problems = self.gate_reply(turn, out)
if problems:
# one repair: the gate's reasons go back to the model; a second refusal escalates
retry = self.chat_json(pb.repair_prompt(prompt, out, problems), tier="default")
again = self.gate_reply(turn, retry)
...
if problems:
return {"reply": pb.GATE_REPLY, "fact_ids": [], "source": "gate", "gated": problems}
After these fixes, twelve varied opening messages were run through the local stack, and none was escalated. That is a spot check on the author's machine, not an evaluation. The remaining gaps are in head quality, and section 11 covers them.
8. Visitor context, under consent#
The agent should know who it is talking to before it asks. The widget reads what the site already knows from its own first-party analytics, and only with consent: the channel, the campaign, the landing page, the pages read, a product hint and whether this is a known account. Without consent, the widget sends no context at all, whatever else the message contains.
The worker treats this context as a hypothesis:
- It shapes the opening and the first question. Someone who landed on a product page is asked about that product.
- It never fills a slot. Only the visitor's own words do that.
- The reply may mention the landing page, or that the visitor is returning. It never mentions the campaign, the channel or how many pages they read.
VISITOR CONTEXT: none (no analytics consent). Do not guess who they are.
9. The booking seam#
For book_demo and hand_to_engineer, the next step is a time on the team's shared calendar. Booking belongs to the conversation. The service lists availability for the conversation, the visitor picks a slot, and the service books it. A slot id is derived from the host and the UTC start time, so a slot listed in any time zone books the same instant. While a visitor is completing a booking, a ten-minute lease holds the slot, and a second visitor gets a typed refusal instead of a double booking. The booking record carries the qualification brief, so the host starts the meeting knowing what was learned. Only a name and a work email are collected.
The calendar connector owns free/busy lookups and the calendar event. The worker never proposes times of its own and never records a booking the calendar did not confirm.
Booking in Preview
The calendar delegation needed for live bookings is not set up yet. A mode that disables booking is in review. With booking disabled, the worker still decides book_demo or hand_to_engineer and records it. The visitor is offered the contact path instead of times, and availability and booking requests are refused with a typed error before any lease is written. Live booking is Planned until the delegation is in place.
10. Parity-first engineering#
The assistant started as a single prototype page. The page held the heads, the playbook, retrieval, the gates and the booking card, and it used a hosted model for wording. Iterating there was fast, and it gave the product owner something to talk to on the first day.
The production worker is a Python port of that page, and the page remains the parity oracle. A test harness runs the page's own decision script in a JavaScript runtime and runs the Python worker on the same conversation states and the same scripted model outputs. The two must match turn by turn: intent, sentiment, stage, legal moves and their rule, the chosen move and why, the outcome, the prompt text, the reply, the citations and the gate reasons.
- Change the page first
Every behavior change, including the "punts too often" fixes, was made in the prototype and tried there.
- Rebuild the page
The page is rebuilt from its template, the head artifacts and the knowledge-base export. A byte-exact rebuild test keeps that reproducible.
- Port, and let parity tell you when you are done
The parity test fails until the Python worker makes the same decisions. Small differences between the languages are pinned explicitly. For example, the sentiment level uses half-up rounding, as the page does, not banker's rounding.
Keep the prototype in a repository
The prototype nearly got lost when the temporary workspace that held it was cleaned. It was recovered from its published copy and committed as the reference. Anything the port will be held to belongs in version control on the day it is written.
11. Evaluation and learning#
Be clear about what the current heads are. They were distilled from synthetic, teacher-authored seed conversations. They prove the pipeline end to end: training, admission, pinning, scoring, parity. They are not evidence of selling ability. Known gaps include a confident not_interested read on a skeptical but interested question, no intent for "show me a demo", and a use case that is rarely extracted, which stalls the move to propose.
The path to trusted heads uses the platform's standard learning loop:
- Real labels Planned: labels come from reviewed transcripts of real conversations. Only labels two independent labelling passes agree on are kept, and a sales lead reviews them.
- Calibration against outcomes Planned: a head's calibration is measured against what actually happened in the conversation, never against the labeller.
- Shadow before promotion Planned: a retrained head runs alongside the pinned one and records what it would have done. It is promoted, as a new digest pin, only when the evaluation gate passes.
- The unanswered queue Preview: every gated or knowledge-base-miss turn is recorded with the question, intent, product, move and reason. This is the list of facts the knowledge base should gain.
The per-turn record is what makes this possible. Every row names the head digests that produced it, so a later outcome can be traced to the exact heads and rules that were running.
12. What comes next#
The implementability gate Planned#
Today, a question outside what the knowledge base covers is caught late: retrieval misses, or the gate refuses the reply. The next piece decides earlier. Before any wording, the request is lowered against the ontologies of the loaded packs: the playbook's actions, the product and capability ontology, and the knowledge base's topic index. The result is one of three verdicts:
| Verdict | What the turn does |
|---|---|
implementable | The heads and playbook choose the move as today, and retrieval must return facts for the covered concepts. |
needs_input | The move is to ask for the missing input, which is usually a qualification slot. |
out_of_scope | An honest, absolute statement that the request is outside what Helixor does, or a hand-off. No model invention. |
The gate is deterministic and fail-closed. If lowering fails, the verdict is out_of_scope with a diagnostic, never an implicit yes. Coverage comes only from declared pack ontologies, never from special-cased words. The verdict and the covered elements will be added to the turn record. The sales assistant is the first surface for the gate, and it will follow the same parity workflow: page first, then port.
A coverage-gap backlog Planned#
out_of_scope and needs_input verdicts will be recorded as a backlog of concepts, missing shapes, surfaces and counts, with no payload text from public surfaces. The backlog decides which ontology concepts and packs to add next, and it will be visible to the sales lead and platform owners.
What to take from this build#
- Put every choice in something you can pin, test and replay: a head, a rule or a gate. Give the language model the words only.
- Let the neural parts propose and let declared rules admit. This applies to slots (verbatim spans), moves (legal set first) and replies (gate first).
- Scale how confident a read must be to what acting on it would do.
- Hand the model short keys, never long identifiers, and map them back.
- Build a fast prototype, keep it as the oracle, and port against it turn by turn.