Skip to content
Firewall latency benchmark

Firewall p95: 2.57msreproducible, on real hardware.

Published numbers, reproducible methodology, no marketing claims without measurements behind them.

Headline result

2.57 ms p95 across 20,000 runs— all four detection layers on, single-threaded, on real hardware. Full methodology below.

p50
2.11 ms
p95
2.57 ms
SLA target 50ms
p99
3.19 ms
max
10.58 ms
worst single request
mean
2.01 ms
throughput
498 req/s
single thread
total runs
20,000

Where the time goes

Per-layer breakdown

Each request runs through up to four layers. The slowest single layer dominates the budget.

Layer
What it does
p50
p95
p99
pattern
Regex-based prompt injection / DAN / jailbreak / system-override patterns. Compiled at boot, no per-call overhead.
0.11
0.16
0.25
token
DLP token scanning — 235 patterns covering SSN, credit card, AWS keys, API keys, etc. Linear scan, short-circuited on first hit.
0.02
0.03
0.04
semantic
Semantic similarity scoring against an embedding-based attack corpus. Dominates the latency budget — single ML inference.
1.94
2.36
2.94
output
Output-side scan only runs after the model responds, so it doesn't add to the request-path latency reported here.
0.00
0.00
0.00

How we measured

Methodology

  • 10 sample promptscovering benign queries, prompt injection ("ignore all previous instructions"), DAN jailbreaks, PII (SSN), system-prompt-leak attempts, base64-encoded payloads, and translations.
  • 200-iteration JIT warm-upbefore measurement so the V8 hot path doesn't poison the p50.
  • Single thread, no concurrency. Throughput under concurrency would be higher; we publish the conservative number.
  • All four enabled layers: pattern + token + semantic + output. The semantic layer alone takes ~95% of the budget — pattern, token, output are sub-millisecond.
  • The public benchmark runs the production artifact. The reproduction below imports the published@evalguard/core npm bundle — the same compiled JS production executes — so there is no source-vs-bundle gap in what you measure.

Run it yourself

Reproduce this on your hardware

Clone the public repo, run one command, get the same JSON shape.

# Public benchmark mirror — runs the published @evalguard/core engine
git clone https://github.com/EvalGuardAi/evalguard-benchmarks
cd evalguard-benchmarks
npm install

# Default run — 5,000 iterations, plain stdout
node benchmark-firewall-latency.mjs

# JSON output (same shape this page reads)
node benchmark-firewall-latency.mjs --runs=20000 --json > latency.json

# Or via the npm script
npm run bench

The mirror's own public CI re-runs this weekly and on every push, failing past 50ms p95 — see the runs.

Chain of custody

Provenance

Measured at
April 30, 2026 · 04:53:59 GMT UTC
Source
scripts/benchmark-firewall-latency.mjs

Numbers are refreshed on every release that touches the firewall path. CI gate prevents regressions.

Want full SLA + uptime guarantees?

The latency budget here is a code-level measurement. Per-tier SLA + uptime history lives on the SLA page.

View SLA