Run decisions over a batch with an audit trail
You run every row of a CSV export through a decision before it leaves your system. Each row is kept, remedied or held back according to the pack's actions, and every row gets one audit line that records what was decided without recording the data. With the built-in example pack (data protection), rows with contact details are written out redacted and rows with fatal data are held back.
About the example
This tutorial uses the built-in data-protection pack so you can run everything without writing a pack first. The same calls work for any decision pack: eligibility, limits, routing and so on. See Core concepts.
What you'll build#
A command, redact_csv.py <in.csv> <clean.csv> <audit.jsonl>, that produces:
clean.csv: rows that were clean or could be repaired, with contact details redacted.audit.jsonl: one JSON line per input row with the row key, the outcome, the pack ID and version, and for each field that fired, its action, rule IDs andreceipt_hash. It never contains field values ormatched_items.- A summary with counts per outcome and per rule.
Prerequisites#
- Make your first decision.
- Read Batch evaluation for throughput guidance.
Steps#
- Create a sample input
Five customer rows: one clean except for contact details, three with a fatal value in
notes, and one with no sensitive data, so the example pack produces every outcome.customer_id,name,email,phone,notes CUST-001,Alice Johnson,alice@example.com,(415) 555-0101,Enterprise renewal Q4 CUST-002,Bob Smith,bob.smith@example.com,(212) 555-0199,SSN on file: 123-45-6789 CUST-003,Carol Davis,carol@example.com,(650) 555-0142,Card ending 4111-1111-1111-1111 CUST-004,Dave Wilson,,,Standard account CUST-005,Eve Torres,eve.torres@example.org,(617) 555-0133,Patient MRN-1234567 referral
- Decide the policy per row
Decide before you write code what each outcome means for a row. With the example pack, this tutorial uses:
Worst field result Row outcome Written to clean.csv?All fields permit cleanYes, unchanged Any field redacts, none blocks redactedYes, with remedy.clean_textin the affected fieldsAny field blocks quarantinedNo. Only its audit line is written. A fatal action means the value must not leave your system, even remedied, so a quarantined row is not written at all. Route quarantined keys to whoever owns the source data.
- Write the pipeline
Evaluate each text field on its own, so the audit line can say which field fired. The engine is created once and reused for every field.
import csv import json import sys import time from collections import Counter from helixor_runtime import HelixorEngine TEXT_FIELDS = ["name", "email", "phone", "notes"] KEY_FIELD = "customer_id" engine = HelixorEngine() def process(src_path: str, clean_path: str, audit_path: str) -> Counter: counts: Counter = Counter() t0 = time.perf_counter() with open(src_path, newline="") as src, \ open(clean_path, "w", newline="") as clean, \ open(audit_path, "w") as audit: reader = csv.DictReader(src) writer = csv.DictWriter(clean, fieldnames=reader.fieldnames) writer.writeheader() for row in reader: fields = {} out = dict(row) blocked = False for name in TEXT_FIELDS: value = row.get(name) or "" if not value: continue d = engine.evaluate(value) if d.invariants_passed: continue fields[name] = { "action": d.action, "rule_ids": [t.rule_id for t in d.triggers], "receipt_hash": d.receipt_hash, } counts.update(t.rule_id for t in d.triggers) if d.action.startswith("block"): blocked = True else: out[name] = d.remedy.clean_text outcome = "quarantined" if blocked else ("redacted" if fields else "clean") counts[outcome] += 1 if not blocked: writer.writerow(out) # One audit line per row. No field values, no matched_items. audit.write(json.dumps({ "key": row[KEY_FIELD], "outcome": outcome, "pack_id": engine.pack_id, "pack_version": engine.version, "fields": fields, }) + "\n") counts["rows"] = sum(counts[k] for k in ("clean", "redacted", "quarantined")) counts["elapsed_ms"] = round((time.perf_counter() - t0) * 1000, 1) return counts if __name__ == "__main__": src, clean, audit = sys.argv[1:4] c = process(src, clean, audit) print(f"rows={c['rows']} clean={c['clean']} redacted={c['redacted']} " f"quarantined={c['quarantined']} elapsed_ms={c['elapsed_ms']}") for key, n in sorted(c.items()): if key.startswith("RULE-"): print(f" {key:26} {n}") - Run it
python redact_csv.py customers.csv clean.csv audit.jsonl
rows=5 clean=1 redacted=1 quarantined=3 elapsed_ms=1.1 RULE-GDPR-EMAIL-REDACT 4 RULE-GLBA-SSN-BLOCK 1 RULE-HIPAA-PHI-BLOCK 1 RULE-PCI-DSS-PAN-BLOCK 1 RULE-TCPA-PHONE-REDACT 4
elapsed_msis wall-clock time for the whole file and varies by machine.clean.csvholds the two rows that may leave:customer_id,name,email,phone,notes CUST-001,Alice Johnson,[REDACTED_EMAIL],[REDACTED_PHONE],Enterprise renewal Q4 CUST-004,Dave Wilson,,,Standard account
Each line of
audit.jsonlis one row. The quarantined rowCUST-002, formatted for reading:{ "key": "CUST-002", "outcome": "quarantined", "pack_id": "compliance.regulatory_pii_guard.v1", "pack_version": "1.0.0", "fields": { "email": { "action": "redact_and_permit_contact_pii", "rule_ids": ["RULE-GDPR-EMAIL-REDACT"], "receipt_hash": "hx_proof_f8a4e44b52d31f36b16ebbf1" }, "phone": { "action": "redact_and_permit_contact_pii", "rule_ids": ["RULE-TCPA-PHONE-REDACT"], "receipt_hash": "hx_proof_42724568e59dfe11bf6a0fa4" }, "notes": { "action": "block_glba_ssn_leakage", "rule_ids": ["RULE-GLBA-SSN-BLOCK"], "receipt_hash": "hx_proof_614a01331f3d28650b8d9eaa" } } }The clean row
CUST-004gets a line too, with"outcome": "clean"and emptyfields, so the audit log accounts for every input row.
How it works#
Each field is a separate decision with its own receipt. A receipt is a fingerprint of the pack, the action, the rules that fired and a SHA-256 of the field value, so the audit log can later show what was decided for a specific value without storing it: given the original value, recompute the receipt and compare. Receipts carry no timestamp, so the same value always gives the same receipt. See Receipts.
Recording pack_id and pack_version on every line matters when the policy changes: an old audit line then still says which rules it was judged by.
The summary counts rules across fields, which is why RULE-GDPR-EMAIL-REDACT shows 4 although only one row was written redacted: the other three emails were in quarantined rows.
When the key is itself sensitive
This tutorial logs customer_id as the row key. If your key is a personal identifier, log a keyed hash of it instead, and keep the hashing key with the source data.
Scaling up#
- Large files. The script streams rows, so memory stays flat. For throughput numbers on your hardware, run the benchmark scripts;
latency_uscovers evaluation only. - Parallelism. Split the file and run one process per core, each with its own engine and its own output files, then concatenate. See Performance.
- Re-runs. Because receipts are deterministic, a re-run over the same input produces the same audit lines, which makes diffs between runs meaningful.