helixordevelopers

Run decisions over a batch with an audit trail

You run every row of a CSV export through a decision before it leaves your system. Each row is kept, remedied or held back according to the pack's actions, and every row gets one audit line that records what was decided without recording the data. With the built-in example pack (data protection), rows with contact details are written out redacted and rows with fatal data are held back.

Tier 025 minutesIntermediatePython 3.9+

About the example

This tutorial uses the built-in data-protection pack so you can run everything without writing a pack first. The same calls work for any decision pack: eligibility, limits, routing and so on. See Core concepts.

What you'll build#

A command, redact_csv.py <in.csv> <clean.csv> <audit.jsonl>, that produces:

  • clean.csv: rows that were clean or could be repaired, with contact details redacted.
  • audit.jsonl: one JSON line per input row with the row key, the outcome, the pack ID and version, and for each field that fired, its action, rule IDs and receipt_hash. It never contains field values or matched_items.
  • A summary with counts per outcome and per rule.

Prerequisites#

Steps#

  1. Create a sample input

    Five customer rows: one clean except for contact details, three with a fatal value in notes, and one with no sensitive data, so the example pack produces every outcome.

    customer_id,name,email,phone,notes
    CUST-001,Alice Johnson,alice@example.com,(415) 555-0101,Enterprise renewal Q4
    CUST-002,Bob Smith,bob.smith@example.com,(212) 555-0199,SSN on file: 123-45-6789
    CUST-003,Carol Davis,carol@example.com,(650) 555-0142,Card ending 4111-1111-1111-1111
    CUST-004,Dave Wilson,,,Standard account
    CUST-005,Eve Torres,eve.torres@example.org,(617) 555-0133,Patient MRN-1234567 referral
    
  2. Decide the policy per row

    Decide before you write code what each outcome means for a row. With the example pack, this tutorial uses:

    Worst field resultRow outcomeWritten to clean.csv?
    All fields permitcleanYes, unchanged
    Any field redacts, none blocksredactedYes, with remedy.clean_text in the affected fields
    Any field blocksquarantinedNo. Only its audit line is written.

    A fatal action means the value must not leave your system, even remedied, so a quarantined row is not written at all. Route quarantined keys to whoever owns the source data.

  3. Write the pipeline

    Evaluate each text field on its own, so the audit line can say which field fired. The engine is created once and reused for every field.

    import csv
    import json
    import sys
    import time
    from collections import Counter
    
    from helixor_runtime import HelixorEngine
    
    TEXT_FIELDS = ["name", "email", "phone", "notes"]
    KEY_FIELD = "customer_id"
    
    engine = HelixorEngine()
    
    
    def process(src_path: str, clean_path: str, audit_path: str) -> Counter:
        counts: Counter = Counter()
        t0 = time.perf_counter()
    
        with open(src_path, newline="") as src, \
             open(clean_path, "w", newline="") as clean, \
             open(audit_path, "w") as audit:
            reader = csv.DictReader(src)
            writer = csv.DictWriter(clean, fieldnames=reader.fieldnames)
            writer.writeheader()
    
            for row in reader:
                fields = {}
                out = dict(row)
                blocked = False
    
                for name in TEXT_FIELDS:
                    value = row.get(name) or ""
                    if not value:
                        continue
                    d = engine.evaluate(value)
                    if d.invariants_passed:
                        continue
                    fields[name] = {
                        "action": d.action,
                        "rule_ids": [t.rule_id for t in d.triggers],
                        "receipt_hash": d.receipt_hash,
                    }
                    counts.update(t.rule_id for t in d.triggers)
                    if d.action.startswith("block"):
                        blocked = True
                    else:
                        out[name] = d.remedy.clean_text
    
                outcome = "quarantined" if blocked else ("redacted" if fields else "clean")
                counts[outcome] += 1
                if not blocked:
                    writer.writerow(out)
    
                # One audit line per row. No field values, no matched_items.
                audit.write(json.dumps({
                    "key": row[KEY_FIELD],
                    "outcome": outcome,
                    "pack_id": engine.pack_id,
                    "pack_version": engine.version,
                    "fields": fields,
                }) + "\n")
    
        counts["rows"] = sum(counts[k] for k in ("clean", "redacted", "quarantined"))
        counts["elapsed_ms"] = round((time.perf_counter() - t0) * 1000, 1)
        return counts
    
    
    if __name__ == "__main__":
        src, clean, audit = sys.argv[1:4]
        c = process(src, clean, audit)
        print(f"rows={c['rows']} clean={c['clean']} redacted={c['redacted']} "
              f"quarantined={c['quarantined']} elapsed_ms={c['elapsed_ms']}")
        for key, n in sorted(c.items()):
            if key.startswith("RULE-"):
                print(f"  {key:26} {n}")
    
  4. Run it
    python redact_csv.py customers.csv clean.csv audit.jsonl
    
    rows=5 clean=1 redacted=1 quarantined=3 elapsed_ms=1.1
      RULE-GDPR-EMAIL-REDACT     4
      RULE-GLBA-SSN-BLOCK        1
      RULE-HIPAA-PHI-BLOCK       1
      RULE-PCI-DSS-PAN-BLOCK     1
      RULE-TCPA-PHONE-REDACT     4
    

    elapsed_ms is wall-clock time for the whole file and varies by machine. clean.csv holds the two rows that may leave:

    customer_id,name,email,phone,notes
    CUST-001,Alice Johnson,[REDACTED_EMAIL],[REDACTED_PHONE],Enterprise renewal Q4
    CUST-004,Dave Wilson,,,Standard account
    

    Each line of audit.jsonl is one row. The quarantined row CUST-002, formatted for reading:

    {
      "key": "CUST-002",
      "outcome": "quarantined",
      "pack_id": "compliance.regulatory_pii_guard.v1",
      "pack_version": "1.0.0",
      "fields": {
        "email": {
          "action": "redact_and_permit_contact_pii",
          "rule_ids": ["RULE-GDPR-EMAIL-REDACT"],
          "receipt_hash": "hx_proof_f8a4e44b52d31f36b16ebbf1"
        },
        "phone": {
          "action": "redact_and_permit_contact_pii",
          "rule_ids": ["RULE-TCPA-PHONE-REDACT"],
          "receipt_hash": "hx_proof_42724568e59dfe11bf6a0fa4"
        },
        "notes": {
          "action": "block_glba_ssn_leakage",
          "rule_ids": ["RULE-GLBA-SSN-BLOCK"],
          "receipt_hash": "hx_proof_614a01331f3d28650b8d9eaa"
        }
      }
    }
    

    The clean row CUST-004 gets a line too, with "outcome": "clean" and empty fields, so the audit log accounts for every input row.

How it works#

Each field is a separate decision with its own receipt. A receipt is a fingerprint of the pack, the action, the rules that fired and a SHA-256 of the field value, so the audit log can later show what was decided for a specific value without storing it: given the original value, recompute the receipt and compare. Receipts carry no timestamp, so the same value always gives the same receipt. See Receipts.

Recording pack_id and pack_version on every line matters when the policy changes: an old audit line then still says which rules it was judged by.

The summary counts rules across fields, which is why RULE-GDPR-EMAIL-REDACT shows 4 although only one row was written redacted: the other three emails were in quarantined rows.

When the key is itself sensitive

This tutorial logs customer_id as the row key. If your key is a personal identifier, log a keyed hash of it instead, and keep the hashing key with the source data.

Scaling up#

  • Large files. The script streams rows, so memory stays flat. For throughput numbers on your hardware, run the benchmark scripts; latency_us covers evaluation only.
  • Parallelism. Split the file and run one process per core, each with its own engine and its own output files, then concatenate. See Performance.
  • Re-runs. Because receipts are deterministic, a re-run over the same input produces the same audit lines, which makes diffs between runs meaningful.

Next steps#