Skip to content
Full Platform

The widest eval + attack coverage in the category. One control plane. 

245 scorers, 345 attack plugins, 91 providers — every number from the drift-checked registry. Six products behind one API key and one dashboard: evaluation, red teaming, firewall and gateway, observability and FinOps, compliance, and agents.

SOC 2 evidence engineISO 42001 mappedEU AI ActGDPR
app.evalguard.ai/dashboard/evals/scorers
EvalGuard Scorer Registry — 245 grading functions across 12 categories with live search and category filters
245
Eval scorers
345
Attack plugins
91
Providers
50
Compliance frameworks

Proof, not promises

The widest coverage in the category — measured.

Our numbers come from the drift-checked registry and the published firewall benchmark. Competitor figures are hand-recorded from their public docs, the coverage donuts are control-map counts, and the load curve is a projection — each is labelled below. Hover any bar to read the count.

Eval scorers vs the field

Built-in, production-ready scorers — no custom authoring required.

EvalGuard245
DeepEval45
MLflow12
Langfuse10

built-in scorers

EvalGuard’s figure is the drift-checked registry. The three competitor counts are theirs, read from their public documentation on 2026-07-22 — the same verification pass /compare dates. They move when those vendors ship; we do not re-measure them for you.

Red-team plugins vs the field

Attack coverage across 100+ strategies and 30 categories.

EvalGuard345
Promptfoo125
Garak37
PyRIT15

attack plugins

EvalGuard’s figure is the drift-checked registry. Promptfoo, Garak and PyRIT are theirs, read from their public documentation on 2026-07-22; Garak publishes its probe count as “37+”, plotted here at its floor. Of our 102 strategies, 15 are adaptive — see the red-team chapter below for what that engine actually does.

Why p95 shouldn’t climb with load

Request path, all layers that run on it. The measured point is 3.67ms p95 over 20,000 runs; the curve is a projection.

See the benchmark →
039771161541005001K5K10K25Kms p95
EvalGuard firewall (~3.67ms p95)Naive in-band scanner

Illustrative. The measurement behind it is one run: 20,000 sequential requests at ~388 req/s single-thread on 2026-08-10, published with its methodology at /trust/latency. Points beyond that throughput are projected from the pre-filter design, and the comparison curve is a generic in-band scanner, not a named product.

Platform breadth

One platform across the full lifecycle — where point tools cover a slice.

EvalRed-teamFirewallGatewayObservabilityCompliance
EvalGuardTypical point tool
100%
OWASP LLM Top 10
10 / 10 mapped
87.5%
OWASP Agentic AI
14 / 16 ASI controls mapped

01 · Eval

Describe your app in one sentence. 245 scorers take it from there.

The NL→Eval pipeline reads a plain-English description of your AI application, maps domain-specific risks, generates targeted test cases, and assembles a production-ready evaluation config. Healthcare, finance, legal: the risk mapping is compliance-aware from the first run.

  • 34 benchmark suites plus custom LLM-as-judge and deterministic assertions
  • 91 providers behind one interface, with a 121-model catalog switchable in one line of config
  • Side-by-side model comparison with regression tracking over time
  • Multi-model orchestration across 77 LLM providers

Also ships: a prompt registry — versioned prompts with side-by-side diffs, one-click rollback, and team approval workflows.

Open the full playground
$“a support agent that reads orders”
generated eval suite
faithfulnesspii-leakjailbreak+18 cases

Live scorer · no signup

Real eval, not a mockup — same deep grader the production API runs.

0/4,000
10 evals / 15 min per IP · no data storedFull playground →

02 · Red Teaming

345 attack plugins. 102 replayable strategies.

The catalogue is deterministic on purpose: every strategy is a transform you can re-run and diff. Layered on top, an opt-in multi-turn adversary probes the model across up to 10 conversation turns (15 max), reads its resistance, and reroutes with UCB1 bandit selection over its own strategy pool. Static test sets miss what an adaptive attacker finds.

  • 102 strategies across 30 attack categories
  • Prompt injection, jailbreak, PII and data-exfiltration probes
  • Parallel attack sessions, each seeded from what already worked in earlier rounds
  • Resistance profiling that re-aims the run at the weakest categories mid-scan

Also ships: a model-file scanner for pickle and safetensors — malicious opcodes and tampered tensors caught before the file ever loads, with ONNX and GGUF structurally audited alongside. Pure TypeScript; attacker code never executes.

Read the attack methodology
adaptive · multi-turnUCB1 bandit
Turn 1
resisted
Turn 2
resisted
Turn 3
breached

03 · Firewall + Gateway

Request-path firewall clears in 3.67ms at p95.

Change one base URL and every LLM call routes through the gateway: input/output firewall, per-key rate limits, cost metering, and SSRF protection across 15 proxied providers. The latency figure is measured — 20,000 runs on 2026-08-10, on the request path — not claimed. The response-side scan is a second call whose cost scales with the response body; it is published separately, by size, at /trust/latency.

  • Visual rule builder with semantic matching, regex patterns, and PII redaction
  • Injection and PII screening on both prompts and responses
  • Streaming support for all 15 proxied providers
  • Reproducible benchmark: pnpm bench:firewall-latency --runs=20000

Also ships: Shadow AI detection (211 AI applications recognised, 20 of them governed by your own approve/restrict/block policy, plus credential-leak flags on outbound prompts), AI-SPM posture scoring with 12 misconfiguration checks, and agentless AI-BOM discovery against your own AWS Bedrock, GCP Vertex AI, and Azure OpenAI credentials.

Read the benchmark
gateway · live1 API
requestroute gpt-4o→claudecache servedfirewall clean200 · 41ms
bench · 20,000 runs
$ pnpm bench:firewall-latency --runs=20000 --write
request path · p95 by layerpattern 0.51mssignature 0.11mstoken 0.04mssemantic 2.86msoutput not measured
p50 2.44msp95 3.67msp99 4.55ms
response scan · p95 by size512 B 2.03ms4 KB 16.52ms96 KB 383.40ms
~388 req/s single-thread · prompts 16–63 B · measured 2026-08-10 · published at /trust/latency

04 · Observability + FinOps

Every trace, token, and dollar, accounted per model.

OpenTelemetry-native ingestion turns agent runs into inspectable trajectories, and every request carries its own cost. Budgets, SLA targets, and drift alerts run on the same stream, so spend and quality never drift apart unnoticed.

  • OTLP trace ingestion with cost breakdown per trace and per model
  • Spend tracking per model, prompt, and team, with budget alerts and forecasting
  • SLA targets you set and track on availability, latency P95/P99, error rate, and throughput
  • Production drift monitoring that alerts on quality regression, with cooldowns and daily limits against alert fatigue

Also ships: a CISO-level risk dashboard — compliance scores, active incidents, and vendor risk on one screen.

See the tracing docs
security posture · live
OWASP LLM Top 1010 / 10
Risk posturelow
Open findings0
99.97%
measured · 90d
41ms
gateway p50
1h
P1 response
0
open incidents

05 · Compliance + Governance

50 frameworks, mapped to evidence instead of screenshots.

Automated assessment across EU AI Act, NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, ISO 42001, India DPDP, HIPAA, and more. Risk classification, gap analysis, and audit-ready documentation generate from your actual runs, not from a questionnaire.

  • EU AI Act risk classification with gap analysis and remediation plans
  • Structured incident response: RCA templates for 11 AI failure categories, auto severity classification
  • Vendor risk: 6-dimension assessments, DPA expiry alerts, lock-in scoring, pre-built profiles for OpenAI, Anthropic, Google, and Mistral
  • A curated register of regulatory changes, plus a print-ready HTML report auditors can save as PDF
Read the compliance methodology
evidence · live12 / 12 CC controls
CC6.1Logical accessevidence
CC7.2Threat detectionevidence
CC8.1Change mgmtevidence

06 · Agents + Voice

Agents traced thought by thought. Voice included.

Distributed tracing for multi-step agents: every thought, action, and observation is a span with its own token count and cost. Voice agents run through the same pipeline, with word-level transcription and audio-deepfake detection instead of a separate voice stack.

  • Agent trajectory visualization with RAG diagnostics and chunk attribution
  • MCP traffic inspection: every tool call risk-scored in real time by the argument firewall
  • Autonomy ladder L1–L4 with firewall-guarded deploys
  • Voice-specific guardrails and latency evals on the realtime audio pipeline

Also ships: a results copilot that prioritizes findings by risk and drafts step-by-step fixes with code examples — today it runs on sample findings; wiring it to your own scans is in progress.

Build an agent on the canvas
otlp · agent trace
thoughtplan refund for order #4821312 tok
actiontool:get_order(#4821)200 · 38ms
observeorder found · refund eligible
voiceasr “refund my last order” · deepfake checkpass
guardautonomy L2 · refund paused for approvalheld

Platform fundamentals

OpenTelemetry export · BYOK encryption · CI/CD templates · self-hosted Docker · Python SDK · OpenAPI docs

Run it against your own model.

The free tier includes 50,000 traces per month. Or get a walkthrough from the team.