helixordevelopers

Methodology

A benchmark number is only useful if you know what it measured and can get it yourself. Every result in this section follows the rules below, and each one links to the script that produced it.

Reporting rules#

  1. Reproducible or not published. Every number comes from a script you can run, listed on Reproduce the results, against artifacts you can obtain.
  2. Hardware and software stated. CPU, core count, memory, operating system, Python and runtime version, and the date of the run.
  3. Absolute, never comparative. We report what Helixor does. We do not publish comparisons with other products.
  4. Distributions, not averages. Latency is reported as p50, p90, p99 and maximum. Means hide the tail that matters in production.
  5. Repeated runs. Each figure is a median of at least three runs, with the spread shown.
  6. Raw output kept. The full script output is published beside the summary.

What each latency metric covers#

The runtime reports its own timing, and it is narrower than what your application sees. Results always say which one they use.

MetricStartsStopsExcludes
Engine time (result.latency_us)Extraction beginsRemedy computedLicense check, receipt hashing, building the result object
Call timeBefore evaluate()After it returnsNothing inside the process
Round tripClient sends the HTTP requestClient has parsed the responseNothing; includes serialization and the loopback network
Cold startFresh process, first evaluate()It returnsInterpreter start-up and imports

Size your latency budget from call time (in-process) or round trip (service), never from engine time.

Run setup#

  • One engine per process, created before timing starts.
  • 100 warm-up calls before measurement; cold start is measured separately in a new process.
  • Timing with time.perf_counter_ns().
  • Fixed input sets, published with the scripts, covering each outcome class: clean, redact and block.
  • Machine otherwise idle, on mains power, with no other benchmark running.
  • Single process unless the result says otherwise. Throughput scales with processes, not threads (see Performance).

What results do not tell you#

  • Your payloads differ. Evaluation cost grows with text length and with the number of matches. Measure with a sample of your own traffic.
  • Laptops are not servers. Frequency scaling and thermal limits change results between runs. Use the results to understand shape and order of magnitude, then measure on your target hardware.
  • Speed is not correctness. Latency results say nothing about detection quality. Test that with a golden set (see Testing policies).