helixordevelopers

Testing policies

You pin down what a pack decides with a table of inputs and expected results, add property checks that must hold for every input, and run the suite in CI so that a runtime upgrade or a playbook edit cannot change a decision without a failing test.

Tier 0Tier 1 for pack tests30 minutesIntermediate

About the example

This tutorial uses the built-in data-protection pack so you can run everything without writing a pack first. The same calls work for any decision pack: eligibility, limits, routing and so on. See Core concepts.

What you'll build#

policy-tests/
├── pytest.ini
├── tests/
│   ├── golden.py          # input → expected action and rule IDs
│   └── test_guard.py      # golden, property, latency and streaming tests
└── pack_tests/
    └── test_internal_ids_pack.py   # optional: tests for your compiled pack

Prerequisites#

Steps#

  1. Write the golden set

    Each case is an input, the action you expect and the exact list of rule IDs, in order. Include near misses that must stay clean: they catch patterns that grow too greedy. Keep the file free of real data; use the standard test values.

    import pytest
    
    # pytest.param(input, expected action, expected rule ids, id=...)
    GOLDEN = [
        pytest.param("Move the design review to Thursday.",
                     "permit_clean_payload", [], id="clean"),
        pytest.param("Send the deck to dana.reyes@example.com.",
                     "redact_and_permit_contact_pii", ["RULE-GDPR-EMAIL-REDACT"], id="email"),
        pytest.param("Call the front desk at (415) 555-0100.",
                     "redact_and_permit_contact_pii", ["RULE-TCPA-PHONE-REDACT"], id="phone"),
        pytest.param("Request came from 192.168.10.24 overnight.",
                     "redact_and_permit_contact_pii", ["RULE-GDPR-IP-REDACT"], id="ipv4"),
        pytest.param("Email ops@example.com or call (212) 555-0199.",
                     "redact_and_permit_contact_pii",
                     ["RULE-GDPR-EMAIL-REDACT", "RULE-TCPA-PHONE-REDACT"], id="email+phone"),
        pytest.param("Applicant SSN 123-45-6789 verified.",
                     "block_glba_ssn_leakage", ["RULE-GLBA-SSN-BLOCK"], id="ssn"),
        pytest.param("Card 4111-1111-1111-1111 on file.",
                     "block_pci_dss_pan_leakage", ["RULE-PCI-DSS-PAN-BLOCK"], id="card"),
        pytest.param("Chart MRN-1234567 attached.",
                     "block_hipaa_phi_leakage", ["RULE-HIPAA-PHI-BLOCK"], id="health-id"),
        # Precedence: the first fatal rule decides; every match is reported.
        pytest.param("Card 4111-1111-1111-1111, receipt to jane@example.com.",
                     "block_pci_dss_pan_leakage", ["RULE-PCI-DSS-PAN-BLOCK", "RULE-GDPR-EMAIL-REDACT"],
                     id="card+email"),
        # Near misses that must stay clean.
        pytest.param("Order 4111-1111-1111-1112 shipped.",
                     "permit_clean_payload", [], id="fails-luhn"),
        pytest.param("Extension 555-0100 is the lobby.",
                     "permit_clean_payload", [], id="seven-digit-number"),
    ]
    

    The card+email case documents a behavior worth pinning: the first fatal rule decides the action, and every rule that matched is reported, in evaluation order.

  2. Configure pytest
    [pytest]
    testpaths = tests
    pythonpath = tests
    
  3. Write the tests

    Four kinds of test, each catching a different failure:

    • Golden: the action and rule IDs for each case are exactly as expected.
    • Properties: statements that hold for every input, such as "clean_text never contains a value the pack matched". They catch a remedy regression even for inputs whose action did not change.
    • Latency budget: a loose wall-clock budget. It exists to catch an accidental slowdown by an order of magnitude, such as a pathological regex, not to benchmark. Keep it loose so shared CI runners do not make it flaky.
    • Streaming: the streamed output equals the block result for several chunk sizes, down to one character per chunk.
    import statistics
    import time
    
    import pytest
    
    from helixor_runtime import HelixorEngine
    
    from golden import GOLDEN
    
    
    @pytest.fixture(scope="module")
    def engine():
        return HelixorEngine()
    
    
    @pytest.mark.parametrize("text,action,rule_ids", GOLDEN)
    def test_golden(engine, text, action, rule_ids):
        d = engine.evaluate(text)
        assert d.action == action
        assert [t.rule_id for t in d.triggers] == rule_ids
        assert d.invariants_passed == (action == "permit_clean_payload")
    
    
    @pytest.mark.parametrize("text,action,rule_ids", GOLDEN)
    def test_clean_text_never_contains_a_match(engine, text, action, rule_ids):
        d = engine.evaluate(text)
        for trigger in d.triggers:
            for item in trigger.matched_items:
                assert item not in d.remedy.clean_text
    
    
    def test_permit_returns_input_unchanged(engine):
        text = "Move the design review to Thursday."
        assert engine.evaluate(text).remedy.clean_text == text
    
    
    def test_dict_input_matches_string_input(engine):
        text = "Send the deck to dana.reyes@example.com."
        assert engine.evaluate({"message": text}).action == engine.evaluate(text).action
    
    
    def test_receipts_are_deterministic(engine):
        text = "Applicant SSN 123-45-6789 verified."
        assert engine.evaluate(text).receipt_hash == engine.evaluate(text).receipt_hash
    
    
    def test_embedded_evaluation_has_no_egress(engine):
        d = engine.evaluate("Send the deck to dana.reyes@example.com.")
        assert d.tokens_spent == 0 and d.egress_bytes == 0
    
    
    def test_latency_budget(engine):
        texts = [g.values[0] for g in GOLDEN]
        engine.evaluate(texts[0])  # warm up
        samples = []
        for _ in range(50):
            for text in texts:
                t0 = time.perf_counter()
                engine.evaluate(text)
                samples.append((time.perf_counter() - t0) * 1e6)
        p50 = statistics.median(samples)
        # Loose, wall-clock budget: catches order-of-magnitude regressions,
        # not noise on shared CI runners.
        assert p50 < 2_000, f"p50 {p50:.0f} µs"
    
    
    STREAMED = "Reach support at support@example.com or call (415) 555-0100."
    
    
    def chunks(text, size):
        return [text[i : i + size] for i in range(0, len(text), size)]
    
    
    @pytest.mark.parametrize("size", [1, 2, 3, 5, 8, 64])
    def test_streaming_matches_block_redaction(engine, size):
        streamed = "".join(engine.stream_filter(chunks(STREAMED, size)))
        assert streamed == engine.evaluate(STREAMED).remedy.clean_text
    
  4. Run the suite
    pytest -v
    
    collecting ... collected 33 items
    
    tests/test_guard.py::test_golden[clean] PASSED                           [  3%]
    tests/test_guard.py::test_golden[email] PASSED                           [  6%]
    tests/test_guard.py::test_golden[phone] PASSED                           [  9%]
    tests/test_guard.py::test_golden[ipv4] PASSED                            [ 12%]
    tests/test_guard.py::test_golden[email+phone] PASSED                     [ 15%]
    tests/test_guard.py::test_golden[ssn] PASSED                             [ 18%]
    tests/test_guard.py::test_golden[card] PASSED                            [ 21%]
    tests/test_guard.py::test_golden[health-id] PASSED                       [ 24%]
    tests/test_guard.py::test_golden[card+email] PASSED                      [ 27%]
    tests/test_guard.py::test_golden[fails-luhn] PASSED                      [ 30%]
    tests/test_guard.py::test_golden[seven-digit-number] PASSED              [ 33%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[clean] PASSED [ 36%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[email] PASSED [ 39%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[phone] PASSED [ 42%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[ipv4] PASSED [ 45%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[email+phone] PASSED [ 48%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[ssn] PASSED  [ 51%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[card] PASSED [ 54%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[health-id] PASSED [ 57%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[card+email] PASSED [ 60%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[fails-luhn] PASSED [ 63%]
    tests/test_guard.py::test_clean_text_never_contains_a_match[seven-digit-number] PASSED [ 66%]
    tests/test_guard.py::test_permit_returns_input_unchanged PASSED          [ 69%]
    tests/test_guard.py::test_dict_input_matches_string_input PASSED         [ 72%]
    tests/test_guard.py::test_receipts_are_deterministic PASSED              [ 75%]
    tests/test_guard.py::test_embedded_evaluation_has_no_egress PASSED       [ 78%]
    tests/test_guard.py::test_latency_budget PASSED                          [ 81%]
    tests/test_guard.py::test_streaming_matches_block_redaction[1] PASSED    [ 84%]
    tests/test_guard.py::test_streaming_matches_block_redaction[2] PASSED    [ 87%]
    tests/test_guard.py::test_streaming_matches_block_redaction[3] PASSED    [ 90%]
    tests/test_guard.py::test_streaming_matches_block_redaction[5] PASSED    [ 93%]
    tests/test_guard.py::test_streaming_matches_block_redaction[8] PASSED    [ 96%]
    tests/test_guard.py::test_streaming_matches_block_redaction[64] PASSED   [100%]
    
    ============================== 33 passed in 0.15s ==============================
    

    This is also how you take a runtime upgrade. The version of this suite written for 0.2.0 pinned two 0.2.0 behaviors. Run against 0.2.1 unchanged, it fails exactly there:

    FAILED tests/test_guard.py::test_golden[card+email] - AssertionError: assert ...
    FAILED tests/test_guard.py::test_streaming_single_character_tokens - [XPASS(s...
    2 failed, 27 passed in 0.15s
    

    Both failures are intended changes: 0.2.1 reports every matching rule, and its streaming no longer damages values split across short chunks. Updating the card+email row and replacing the expected failure with the streaming test above is the whole review.

  5. Test a compiled pack

    Load the pack you compiled in Add custom rules once per module with HelixorEngine.load_pack(). Assert pack_id first, so a test can never pass against the wrong pack.

    import os
    
    import pytest
    
    from helixor_runtime import HelixorEngine
    
    PACK = os.environ["POLICY_PACK"]          # e.g. internal_ids.hxpack
    LICENSE = os.environ["POLICY_LICENSE"]    # e.g. ~/.helixor/helixor.lic
    
    
    @pytest.fixture(scope="module")
    def engine():
        engine = HelixorEngine.load_pack(PACK, license_file=LICENSE)
        assert engine.pack_id == "custom.internal_ids.v1"
        return engine
    
    
    @pytest.mark.parametrize("text,action", [
        pytest.param("Status for PROJ-ZEUS-9X is green.", "block_internal_project_code", id="project-code"),
        pytest.param("Kickoff PROJ-APOLLO-12 tomorrow.", "block_internal_project_code", id="other-project"),
        pytest.param("Email ops@example.com about PROJ-ZEUS-9X.", "block_internal_project_code", id="code+email"),
        pytest.param("See TCK-004211 for the fix.", "redact_ticket_id", id="ticket"),
        pytest.param("The PROJ-ZEUS team meets today.", "permit_clean_payload", id="no-number"),
        pytest.param("Lunch is at noon.", "permit_clean_payload", id="clean"),
    ])
    def test_pack_actions(engine, text, action):
        assert engine.evaluate(text).action == action
    
    
    def test_redact_rule_removes_the_ticket(engine):
        assert "TCK-004211" not in engine.evaluate("See TCK-004211 for the fix.").remedy.clean_text
    
    POLICY_PACK=internal_ids.hxpack POLICY_LICENSE="$HOME/.helixor/helixor.lic" pytest -v pack_tests
    
    collecting ... collected 7 items
    
    pack_tests/test_internal_ids_pack.py::test_pack_actions[project-code] PASSED [ 14%]
    pack_tests/test_internal_ids_pack.py::test_pack_actions[other-project] PASSED [ 28%]
    pack_tests/test_internal_ids_pack.py::test_pack_actions[code+email] PASSED [ 42%]
    pack_tests/test_internal_ids_pack.py::test_pack_actions[ticket] PASSED   [ 57%]
    pack_tests/test_internal_ids_pack.py::test_pack_actions[no-number] PASSED [ 71%]
    pack_tests/test_internal_ids_pack.py::test_pack_actions[clean] PASSED    [ 85%]
    pack_tests/test_internal_ids_pack.py::test_redact_rule_removes_the_ticket PASSED [100%]
    
    ============================== 7 passed in 0.12s ===============================
    

    The test fails with a KeyError if either variable is unset, and with FileNotFoundError if either file is missing. That is deliberate: a pack test that silently skips is not a test.

  6. Run it in CI

    Add one step to your pipeline that installs the pinned runtime version and runs the suite. Until helixor-runtime is on the public package index, install the wheel you were given instead; see Installation. A non-zero exit code fails the build; the JUnit XML report gives most CI systems a per-test view.

    python -m pip install "helixor-runtime==0.2.1" pytest
    python -m pytest -q --junitxml=reports/policy.xml
    
    .................................                                        [100%]
    33 passed in 0.15s
    

    Run the pack tests in a separate step that has the license available as a CI secret file. Never commit the license to the repository.

How it works#

Decisions are deterministic: the same pack and the same input always give the same action, rule IDs, clean text and receipt. That makes a golden set a complete specification of the behavior you rely on, and any diff in it is a real change, never noise. Pin the runtime version in CI and upgrade it deliberately, with the golden set as the gate.

When a test fails after an upgrade or a playbook change, decide whether the new behavior is what you want. If it is, update the golden row in the same change, so reviewers see the policy change next to the code change. See Operational excellence for release practice and Troubleshooting for common failures.

Next steps#