Beyond rules: search, strategy and learning
Rules decide over the facts you give them. These three examples show what an engine does when the answer has to be found: it searches ahead and proves a move, it changes technique when the position allows an exact answer, and it picks a strategy for each situation and learns from how episodes end. Every request and response is shown in the API inspector.
What runs here
The games run on the playground service at /api/try/games/*. These are demo endpoints, not the product API: the engines are ports of the game adapters in the Helixor reasoning service, run only on the server, and are tested to return exactly what the reasoning service returns for the same positions and episodes. The badge above shows the source commit the service reports. Nothing on this page is simulated: when the service cannot be reached the page says so and shows no boards.
Examples#
Search with proofs
Drop a disc to start. After each engine move this panel shows what it searched and why it chose the column.
Which capability this shows
Search, and a proof when there is one. For every legal column the engine searches the replies four plies deep, up to 12,000 positions. When every line of play within that horizon ends in a win for it, the column is certified: a proof, checked by the tests of the playground against an independent full-width search, that the engine wins whatever you play. Otherwise the engine plays the column its search preferred and says so: the preference is a softmax over search values, not a calibrated probability.
Why rules alone cannot do this. A rule evaluates conditions over the facts it is given. Choosing a move means generating the replies to each candidate, and the replies to those, and a proof that a move wins has to cover every one of them. Rules can check that a column is legal, and the engine does that before it searches; they cannot produce or search that tree.
Not shown here. The engine does not learn across four-in-a-row games on this page. Beyond four plies it has no proof, so it can lose. See Ask the reasoning service, and handle abstentions for how answers with and without a proof are reported by the reasoning API.
Choosing the technique for the position
Play a highlighted square. After each engine move this panel shows which technique produced it.
Which capability this shows
The engine picks its technique from the position. With more than 12 empty squares the game tree is out of reach, so the engine runs a depth-limited search with a positional evaluation (corners, mobility) and plays its preferred move. From 12 empty squares down it searches to the end of the game. When that search finishes within its 15,000-node budget, the move is certified and the engine reports the exact final disc margin with best play by both sides. In the tests of the playground the exact search finished for every position tried at 8 or fewer empty squares, often at 9 and rarely at 10 or more; when it runs out of budget the move is labelled endgame_budget_exhausted and is not certified.
Why rules alone cannot do this. A rule set applies the same kind of logic everywhere. Here the engine moves from an estimate to an exact answer when the problem becomes small enough, and every move says which one you got. The margins it certifies are checked in the tests against an exhaustive search with no pruning.
Not shown here. No learning across games, and no opening knowledge.
Strategy selection and learning from outcomes
Run an episode. Engine (blue) against an opponent (coral) that always plays the strategy you picked. This panel then shows the strategy the engine ranked first at each tick and what it learned at the end.
Which capability this shows
Strategy selection. At every tick the engine scores seven strategies (take cover, get health, get armor, keep distance, ambush, rush, flank) against its state: health, armor, whether it is reloading, line of sight, nearby pickups and the time left. The tactical terms of each score are fixed; the engine plays the highest total. A fixed pursuit rule can still override where it moves when neither side has seen the other for a while, and the inspector marks those ticks.
Learning from outcomes. After each episode the engine records the outcome against the strategies in use in a belief ledger: a Beta posterior per strategy, where a contradiction weighs three times a confirmation. The next episode adds each strategy's belief to its score, and strategies with enough evidence against them are penalized heavily. This is the same ledger the runtime ships as HelixorBeliefLedger; see Track outcomes and confidence.
How much it helps. The effect is small and consistent. In 2,100 paired episodes in the playground's tests (seven opponent strategies, three blocks of 100 seeds) the engine won 555 without learning and 618 with learning, and each block improved. Against the cover opponent, what it had learned changed at least one of its strategy choices in 22 of 100 episodes. The engine still loses most duels against these opponents. A paired run of 10 to 30 episodes, as above, is noisy: it can show no gain or a loss.
Why rules alone cannot do this. A rule fires on the conditions someone wrote, the same way every time. Here the ranking of the options changes with the evidence of how earlier episodes ended, and the evidence is kept per strategy with its uncertainty.
Not shown here. The arena is a two-dimensional simulator. The learning is outcome memory only: there is no proposal, golden regression gate or admitted policy version here; that loop is in Close the learning loop. What the engine learns lives in your browser tab: each request sends the ledger back, and the service stores nothing.
API inspector#
Every call this page makes, newest first: method, path, headers, the exact JSON body, the response with its status and timing, and what each field means.
Demo endpoints and the product API#
| This page calls | Kind | What it illustrates in the product |
|---|---|---|
POST /api/try/games/connect4/movePOST /api/try/games/othello/move | Playground demo endpoint, anonymous, rate limited | Game goals game.connect4.v1 and game.othello.v1 on the reasoning service's POST /v1/decide Preview. The reasoning service is not publicly deployed and needs a caller credential. The inspector shows the equivalent decide request for the current position. |
POST /api/try/games/arena/duelPOST /api/try/games/arena/compare | Playground demo endpoint | Outcome memory with HelixorBeliefLedger in the runtime: record, belief, and snapshot / restore Next release. The inspector shows the equivalent calls. |
GET /api/try/games | Playground demo endpoint | The catalog of these examples: budgets and the source commit of each engine. |
Known issue in the reasoning service (Preview)
For game goals, POST /v1/decide currently reports a move that has no proof as an answer whose p_correct is the search preference, which is not a calibrated probability, and a certified four-in-a-row win fails with HTTP 500. Treat a game answer as certified only when proof.kind is solver_certificate. The demo endpoints on this page label both cases correctly.
Limits#
- Budgets are fixed on the server. Four in a row: depth 4, 12,000 positions per move. Reversi: depth 4 or an exact endgame search, 15,000 positions per move. Arena: 35 ticks per episode, 30 episodes per arm of a paired run.
- Shared limits. These calls share the playground's per-address rate limit and 8 KB request cap. A request that does not finish in time returns
GAME_TIMEOUT. - Nothing is stored. Positions come from your browser with each request. The arena's learned beliefs come back in each response as a
helixor.belief_ledger.v1snapshot and go back with the next request; closing the tab forgets them. The service logs the game, status and outcome only.