Testing policies
You pin down what a pack decides with a table of inputs and expected results, add property checks that must hold for every input, and run the suite in CI so that a runtime upgrade or a playbook edit cannot change a decision without a failing test.
About the example
This tutorial uses the built-in data-protection pack so you can run everything without writing a pack first. The same calls work for any decision pack: eligibility, limits, routing and so on. See Core concepts.
What you'll build#
policy-tests/
├── pytest.ini
├── tests/
│ ├── golden.py # input → expected action and rule IDs
│ └── test_guard.py # golden, property, latency and streaming tests
└── pack_tests/
└── test_internal_ids_pack.py # optional: tests for your compiled pack
Prerequisites#
- The runtime and
pytestinstalled in the same environment. - Make your first decision. The pack-test step also assumes Add custom rules.
Steps#
- Write the golden set
Each case is an input, the action you expect and the exact list of rule IDs, in order. Include near misses that must stay clean: they catch patterns that grow too greedy. Keep the file free of real data; use the standard test values.
import pytest # pytest.param(input, expected action, expected rule ids, id=...) GOLDEN = [ pytest.param("Move the design review to Thursday.", "permit_clean_payload", [], id="clean"), pytest.param("Send the deck to dana.reyes@example.com.", "redact_and_permit_contact_pii", ["RULE-GDPR-EMAIL-REDACT"], id="email"), pytest.param("Call the front desk at (415) 555-0100.", "redact_and_permit_contact_pii", ["RULE-TCPA-PHONE-REDACT"], id="phone"), pytest.param("Request came from 192.168.10.24 overnight.", "redact_and_permit_contact_pii", ["RULE-GDPR-IP-REDACT"], id="ipv4"), pytest.param("Email ops@example.com or call (212) 555-0199.", "redact_and_permit_contact_pii", ["RULE-GDPR-EMAIL-REDACT", "RULE-TCPA-PHONE-REDACT"], id="email+phone"), pytest.param("Applicant SSN 123-45-6789 verified.", "block_glba_ssn_leakage", ["RULE-GLBA-SSN-BLOCK"], id="ssn"), pytest.param("Card 4111-1111-1111-1111 on file.", "block_pci_dss_pan_leakage", ["RULE-PCI-DSS-PAN-BLOCK"], id="card"), pytest.param("Chart MRN-1234567 attached.", "block_hipaa_phi_leakage", ["RULE-HIPAA-PHI-BLOCK"], id="health-id"), # Precedence: the first fatal rule decides; every match is reported. pytest.param("Card 4111-1111-1111-1111, receipt to jane@example.com.", "block_pci_dss_pan_leakage", ["RULE-PCI-DSS-PAN-BLOCK", "RULE-GDPR-EMAIL-REDACT"], id="card+email"), # Near misses that must stay clean. pytest.param("Order 4111-1111-1111-1112 shipped.", "permit_clean_payload", [], id="fails-luhn"), pytest.param("Extension 555-0100 is the lobby.", "permit_clean_payload", [], id="seven-digit-number"), ]The
card+emailcase documents a behavior worth pinning: the first fatal rule decides the action, and every rule that matched is reported, in evaluation order. - Configure pytest
[pytest] testpaths = tests pythonpath = tests
- Write the tests
Four kinds of test, each catching a different failure:
- Golden: the action and rule IDs for each case are exactly as expected.
- Properties: statements that hold for every input, such as "
clean_textnever contains a value the pack matched". They catch a remedy regression even for inputs whose action did not change. - Latency budget: a loose wall-clock budget. It exists to catch an accidental slowdown by an order of magnitude, such as a pathological regex, not to benchmark. Keep it loose so shared CI runners do not make it flaky.
- Streaming: the streamed output equals the block result for several chunk sizes, down to one character per chunk.
import statistics import time import pytest from helixor_runtime import HelixorEngine from golden import GOLDEN @pytest.fixture(scope="module") def engine(): return HelixorEngine() @pytest.mark.parametrize("text,action,rule_ids", GOLDEN) def test_golden(engine, text, action, rule_ids): d = engine.evaluate(text) assert d.action == action assert [t.rule_id for t in d.triggers] == rule_ids assert d.invariants_passed == (action == "permit_clean_payload") @pytest.mark.parametrize("text,action,rule_ids", GOLDEN) def test_clean_text_never_contains_a_match(engine, text, action, rule_ids): d = engine.evaluate(text) for trigger in d.triggers: for item in trigger.matched_items: assert item not in d.remedy.clean_text def test_permit_returns_input_unchanged(engine): text = "Move the design review to Thursday." assert engine.evaluate(text).remedy.clean_text == text def test_dict_input_matches_string_input(engine): text = "Send the deck to dana.reyes@example.com." assert engine.evaluate({"message": text}).action == engine.evaluate(text).action def test_receipts_are_deterministic(engine): text = "Applicant SSN 123-45-6789 verified." assert engine.evaluate(text).receipt_hash == engine.evaluate(text).receipt_hash def test_embedded_evaluation_has_no_egress(engine): d = engine.evaluate("Send the deck to dana.reyes@example.com.") assert d.tokens_spent == 0 and d.egress_bytes == 0 def test_latency_budget(engine): texts = [g.values[0] for g in GOLDEN] engine.evaluate(texts[0]) # warm up samples = [] for _ in range(50): for text in texts: t0 = time.perf_counter() engine.evaluate(text) samples.append((time.perf_counter() - t0) * 1e6) p50 = statistics.median(samples) # Loose, wall-clock budget: catches order-of-magnitude regressions, # not noise on shared CI runners. assert p50 < 2_000, f"p50 {p50:.0f} µs" STREAMED = "Reach support at support@example.com or call (415) 555-0100." def chunks(text, size): return [text[i : i + size] for i in range(0, len(text), size)] @pytest.mark.parametrize("size", [1, 2, 3, 5, 8, 64]) def test_streaming_matches_block_redaction(engine, size): streamed = "".join(engine.stream_filter(chunks(STREAMED, size))) assert streamed == engine.evaluate(STREAMED).remedy.clean_text - Run the suite
pytest -v
collecting ... collected 33 items tests/test_guard.py::test_golden[clean] PASSED [ 3%] tests/test_guard.py::test_golden[email] PASSED [ 6%] tests/test_guard.py::test_golden[phone] PASSED [ 9%] tests/test_guard.py::test_golden[ipv4] PASSED [ 12%] tests/test_guard.py::test_golden[email+phone] PASSED [ 15%] tests/test_guard.py::test_golden[ssn] PASSED [ 18%] tests/test_guard.py::test_golden[card] PASSED [ 21%] tests/test_guard.py::test_golden[health-id] PASSED [ 24%] tests/test_guard.py::test_golden[card+email] PASSED [ 27%] tests/test_guard.py::test_golden[fails-luhn] PASSED [ 30%] tests/test_guard.py::test_golden[seven-digit-number] PASSED [ 33%] tests/test_guard.py::test_clean_text_never_contains_a_match[clean] PASSED [ 36%] tests/test_guard.py::test_clean_text_never_contains_a_match[email] PASSED [ 39%] tests/test_guard.py::test_clean_text_never_contains_a_match[phone] PASSED [ 42%] tests/test_guard.py::test_clean_text_never_contains_a_match[ipv4] PASSED [ 45%] tests/test_guard.py::test_clean_text_never_contains_a_match[email+phone] PASSED [ 48%] tests/test_guard.py::test_clean_text_never_contains_a_match[ssn] PASSED [ 51%] tests/test_guard.py::test_clean_text_never_contains_a_match[card] PASSED [ 54%] tests/test_guard.py::test_clean_text_never_contains_a_match[health-id] PASSED [ 57%] tests/test_guard.py::test_clean_text_never_contains_a_match[card+email] PASSED [ 60%] tests/test_guard.py::test_clean_text_never_contains_a_match[fails-luhn] PASSED [ 63%] tests/test_guard.py::test_clean_text_never_contains_a_match[seven-digit-number] PASSED [ 66%] tests/test_guard.py::test_permit_returns_input_unchanged PASSED [ 69%] tests/test_guard.py::test_dict_input_matches_string_input PASSED [ 72%] tests/test_guard.py::test_receipts_are_deterministic PASSED [ 75%] tests/test_guard.py::test_embedded_evaluation_has_no_egress PASSED [ 78%] tests/test_guard.py::test_latency_budget PASSED [ 81%] tests/test_guard.py::test_streaming_matches_block_redaction[1] PASSED [ 84%] tests/test_guard.py::test_streaming_matches_block_redaction[2] PASSED [ 87%] tests/test_guard.py::test_streaming_matches_block_redaction[3] PASSED [ 90%] tests/test_guard.py::test_streaming_matches_block_redaction[5] PASSED [ 93%] tests/test_guard.py::test_streaming_matches_block_redaction[8] PASSED [ 96%] tests/test_guard.py::test_streaming_matches_block_redaction[64] PASSED [100%] ============================== 33 passed in 0.15s ==============================
This is also how you take a runtime upgrade. The version of this suite written for 0.2.0 pinned two 0.2.0 behaviors. Run against 0.2.1 unchanged, it fails exactly there:
FAILED tests/test_guard.py::test_golden[card+email] - AssertionError: assert ... FAILED tests/test_guard.py::test_streaming_single_character_tokens - [XPASS(s... 2 failed, 27 passed in 0.15s
Both failures are intended changes: 0.2.1 reports every matching rule, and its streaming no longer damages values split across short chunks. Updating the
card+emailrow and replacing the expected failure with the streaming test above is the whole review. - Test a compiled pack
Load the pack you compiled in Add custom rules once per module with
HelixorEngine.load_pack(). Assertpack_idfirst, so a test can never pass against the wrong pack.import os import pytest from helixor_runtime import HelixorEngine PACK = os.environ["POLICY_PACK"] # e.g. internal_ids.hxpack LICENSE = os.environ["POLICY_LICENSE"] # e.g. ~/.helixor/helixor.lic @pytest.fixture(scope="module") def engine(): engine = HelixorEngine.load_pack(PACK, license_file=LICENSE) assert engine.pack_id == "custom.internal_ids.v1" return engine @pytest.mark.parametrize("text,action", [ pytest.param("Status for PROJ-ZEUS-9X is green.", "block_internal_project_code", id="project-code"), pytest.param("Kickoff PROJ-APOLLO-12 tomorrow.", "block_internal_project_code", id="other-project"), pytest.param("Email ops@example.com about PROJ-ZEUS-9X.", "block_internal_project_code", id="code+email"), pytest.param("See TCK-004211 for the fix.", "redact_ticket_id", id="ticket"), pytest.param("The PROJ-ZEUS team meets today.", "permit_clean_payload", id="no-number"), pytest.param("Lunch is at noon.", "permit_clean_payload", id="clean"), ]) def test_pack_actions(engine, text, action): assert engine.evaluate(text).action == action def test_redact_rule_removes_the_ticket(engine): assert "TCK-004211" not in engine.evaluate("See TCK-004211 for the fix.").remedy.clean_textPOLICY_PACK=internal_ids.hxpack POLICY_LICENSE="$HOME/.helixor/helixor.lic" pytest -v pack_tests
collecting ... collected 7 items pack_tests/test_internal_ids_pack.py::test_pack_actions[project-code] PASSED [ 14%] pack_tests/test_internal_ids_pack.py::test_pack_actions[other-project] PASSED [ 28%] pack_tests/test_internal_ids_pack.py::test_pack_actions[code+email] PASSED [ 42%] pack_tests/test_internal_ids_pack.py::test_pack_actions[ticket] PASSED [ 57%] pack_tests/test_internal_ids_pack.py::test_pack_actions[no-number] PASSED [ 71%] pack_tests/test_internal_ids_pack.py::test_pack_actions[clean] PASSED [ 85%] pack_tests/test_internal_ids_pack.py::test_redact_rule_removes_the_ticket PASSED [100%] ============================== 7 passed in 0.12s ===============================
The test fails with a
KeyErrorif either variable is unset, and withFileNotFoundErrorif either file is missing. That is deliberate: a pack test that silently skips is not a test. - Run it in CI
Add one step to your pipeline that installs the pinned runtime version and runs the suite. Until
helixor-runtimeis on the public package index, install the wheel you were given instead; see Installation. A non-zero exit code fails the build; the JUnit XML report gives most CI systems a per-test view.python -m pip install "helixor-runtime==0.2.1" pytest python -m pytest -q --junitxml=reports/policy.xml
................................. [100%] 33 passed in 0.15s
Run the pack tests in a separate step that has the license available as a CI secret file. Never commit the license to the repository.
How it works#
Decisions are deterministic: the same pack and the same input always give the same action, rule IDs, clean text and receipt. That makes a golden set a complete specification of the behavior you rely on, and any diff in it is a real change, never noise. Pin the runtime version in CI and upgrade it deliberately, with the golden set as the gate.
When a test fails after an upgrade or a playbook change, decide whether the new behavior is what you want. If it is, update the golden row in the same change, so reviewers see the policy change next to the code change. See Operational excellence for release practice and Troubleshooting for common failures.