Skip to content
Full Platform

The widest eval + attack coverage in the category. One control plane. 

245 scorers, 343 attack plugins, 91 providers — every number from the drift-checked registry. Six products behind one API key and one dashboard: evaluation, red teaming, firewall and gateway, observability and FinOps, compliance, and agents.

SOC 2 evidence engineISO 42001 mappedEU AI ActGDPR
app.evalguard.ai/dashboard/evals/scorers
EvalGuard Scorer Registry — 245 grading functions across 12 categories with live search and category filters
245
Eval scorers
343
Attack plugins
91
Providers
50
Compliance frameworks

Proof, not promises

The widest coverage in the category — measured.

Every number is sourced from the drift-checked registry and the public firewall benchmark, not a slide. Hover any bar to read the count.

Eval scorers vs the field

Built-in, production-ready scorers — no custom authoring required.

EvalGuard245
DeepEval45
MLflow12
Langfuse10

built-in scorers

Red-team plugins vs the field

Attack coverage across 100+ strategies and 30 categories.

EvalGuard343
Promptfoo125
Garak37
PyRIT15

attack plugins

Firewall p95 stays flat under load

Full pipeline (pattern · token · semantic · output), measured across 20K runs.

See the benchmark →
039771161541005001K5K10K25Kms p95
EvalGuard firewall (~2.57ms p95)Naive in-band scanner

Platform breadth

One platform across the full lifecycle — where point tools cover a slice.

EvalRed-teamFirewallGatewayObservabilityCompliance
EvalGuardTypical point tool
100%
OWASP LLM Top 10
10 / 10 mapped
100%
OWASP Agentic AI
full control map

01 · Eval

Describe your app in one sentence. 245 scorers take it from there.

The NL→Eval pipeline reads a plain-English description of your AI application, maps domain-specific risks, generates targeted test cases, and assembles a production-ready evaluation config. Healthcare, finance, legal: the risk mapping is compliance-aware from the first run.

  • 34 benchmark suites plus custom LLM-as-judge and deterministic assertions
  • 91 providers behind one interface, with a 121-model catalog switchable in one line of config
  • Side-by-side model comparison with regression tracking over time
  • Multi-model orchestration across 77 LLM providers

Also ships: a prompt registry — versioned prompts with side-by-side diffs, one-click rollback, and team approval workflows.

Open the full playground
$“a support agent that reads orders”
generated eval suite
faithfulnesspii-leakjailbreak+18 cases

Live scorer · no signup

Real eval, not a mockup — same deep grader the production API runs.

0/4,000
10 evals / 15 min per IP · no data storedFull playground →

02 · Red Teaming

343 attack plugins. 100 strategies. An attacker that adapts.

An AI adversary probes your model across up to 10 conversation turns, reads its resistance, and reroutes strategy with UCB1 bandit optimization. Static test sets miss what an adaptive attacker finds.

  • 100 strategies across 30 attack categories
  • Prompt injection, jailbreak, PII and data-exfiltration probes
  • Parallel attack sessions with cross-session memory
  • Real-time resistance profiling per target

Also ships: a model-file scanner for pickle, safetensors, GGUF, and ONNX — malicious opcodes and tampered tensors caught before the file ever loads. Pure TypeScript; attacker code never executes.

Read the attack methodology
adaptive · multi-turnUCB1 bandit
Turn 1
resisted
Turn 2
resisted
Turn 3
breached

03 · Firewall + Gateway

The full firewall pipeline clears in 2.57ms at p95.

Change one base URL and every LLM call routes through the gateway: input/output firewall, per-key rate limits, cost metering, and SSRF protection across 15 proxied providers. The latency number is measured across 20,000 runs, not claimed.

  • Visual rule builder with semantic matching, regex patterns, and PII redaction
  • Injection and PII screening on both prompts and responses
  • Streaming support for all 15 proxied providers
  • Reproducible benchmark: pnpm bench:firewall-latency

Also ships: Shadow AI detection (20 AI providers tracked, credential-leak flags on outbound prompts), AI-SPM posture scoring with 12 misconfiguration checks, and agentless AI-BOM discovery across AWS Bedrock, GCP Vertex AI, and Azure OpenAI.

Read the benchmark
gateway · live1 API
requestroute gpt-4o→claudecache servedfirewall clean200 · 41ms
bench · 20,000 runs
$ pnpm bench:firewall-latency
pattern 0.16mstoken 0.03mssemantic 2.36msoutput 0.00ms
p50 2.11msp95 2.57msp99 3.19ms
~498 req/s single-thread · 0 errors · published at /trust/latency

04 · Observability + FinOps

Every trace, token, and dollar, accounted per model.

OpenTelemetry-native ingestion turns agent runs into inspectable trajectories, and every request carries its own cost. Budgets, SLA targets, and drift alerts run on the same stream, so spend and quality never drift apart unnoticed.

  • OTLP trace ingestion with cost breakdown per trace and per model
  • Spend tracking per model, prompt, and team, with budget alerts and forecasting
  • SLA targets on availability, latency P95/P99, error rate, and throughput, with automatic violation alerts
  • Z-score drift detection triggers re-evaluation automatically, with cooldowns and daily limits against alert fatigue

Also ships: a CISO-level risk dashboard — compliance scores, active incidents with MTTR, and vendor risk on one screen.

See the tracing docs
security posture · live
OWASP LLM Top 1010 / 10
Risk posturelow
Open findings0
99.97%
measured · 90d
41ms
gateway p50
1h
P1 response
0
open incidents

05 · Compliance + Governance

50 frameworks, mapped to evidence instead of screenshots.

Automated assessment across EU AI Act, NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, ISO 42001, India DPDP, HIPAA, and more. Risk classification, gap analysis, and audit-ready documentation generate from your actual runs, not from a questionnaire.

  • EU AI Act risk classification with gap analysis and remediation plans
  • Structured incident response: RCA templates for 11 AI failure categories, auto severity classification
  • Vendor risk: 6-dimension assessments, DPA expiry alerts, lock-in scoring, pre-built profiles for OpenAI, Anthropic, Google, and Mistral
  • Regulatory change tracking with alerts, plus PDF export for auditors
Read the compliance methodology
evidence · live12 / 12 CC controls
CC6.1Logical accessevidence
CC7.2Threat detectionevidence
CC8.1Change mgmtevidence

06 · Agents + Voice

Agents traced thought by thought. Voice included.

Distributed tracing for multi-step agents: every thought, action, and observation is a span with its own token count and cost. Voice agents run through the same pipeline, with word-level transcription and audio-deepfake detection instead of a separate voice stack.

  • Agent trajectory visualization with RAG diagnostics and chunk attribution
  • MCP traffic inspection: every tool call risk-scored in real time, memory poisoning flagged
  • Autonomy ladder L1–L4 with firewall-guarded deploys
  • Voice-specific guardrails and latency evals on the realtime audio pipeline

Also ships: a results copilot that reads scan and eval output, prioritizes findings by risk, and drafts step-by-step fixes with code examples.

Build an agent on the canvas
otlp · agent trace
thoughtplan refund for order #4821312 tok
actiontool:get_order(#4821)200 · 38ms
observeorder found · refund eligible
voiceasr “refund my last order” · deepfake checkpass
guardautonomy L2 · refund paused for approvalheld

Platform fundamentals

OpenTelemetry export · BYOK encryption · CI/CD templates · self-hosted Docker · Python SDK · OpenAPI docs

Run it against your own model.

The free tier includes 50,000 traces per month. Or get a walkthrough from the team.