The widest eval + attack coverage in the category. One control plane.
245 scorers, 343 attack plugins, 91 providers — every number from the drift-checked registry. Six products behind one API key and one dashboard: evaluation, red teaming, firewall and gateway, observability and FinOps, compliance, and agents.

Proof, not promises
The widest coverage in the category — measured.
Every number is sourced from the drift-checked registry and the public firewall benchmark, not a slide. Hover any bar to read the count.
Eval scorers vs the field
Built-in, production-ready scorers — no custom authoring required.
built-in scorers
Red-team plugins vs the field
Attack coverage across 100+ strategies and 30 categories.
attack plugins
Firewall p95 stays flat under load
Full pipeline (pattern · token · semantic · output), measured across 20K runs.
Platform breadth
One platform across the full lifecycle — where point tools cover a slice.
01 · Eval
Describe your app in one sentence. 245 scorers take it from there.
The NL→Eval pipeline reads a plain-English description of your AI application, maps domain-specific risks, generates targeted test cases, and assembles a production-ready evaluation config. Healthcare, finance, legal: the risk mapping is compliance-aware from the first run.
- 34 benchmark suites plus custom LLM-as-judge and deterministic assertions
- 91 providers behind one interface, with a 121-model catalog switchable in one line of config
- Side-by-side model comparison with regression tracking over time
- Multi-model orchestration across 77 LLM providers
Also ships: a prompt registry — versioned prompts with side-by-side diffs, one-click rollback, and team approval workflows.
Open the full playgroundLive scorer · no signup
Real eval, not a mockup — same deep grader the production API runs.
02 · Red Teaming
343 attack plugins. 100 strategies. An attacker that adapts.
An AI adversary probes your model across up to 10 conversation turns, reads its resistance, and reroutes strategy with UCB1 bandit optimization. Static test sets miss what an adaptive attacker finds.
- 100 strategies across 30 attack categories
- Prompt injection, jailbreak, PII and data-exfiltration probes
- Parallel attack sessions with cross-session memory
- Real-time resistance profiling per target
Also ships: a model-file scanner for pickle, safetensors, GGUF, and ONNX — malicious opcodes and tampered tensors caught before the file ever loads. Pure TypeScript; attacker code never executes.
Read the attack methodology03 · Firewall + Gateway
The full firewall pipeline clears in 2.57ms at p95.
Change one base URL and every LLM call routes through the gateway: input/output firewall, per-key rate limits, cost metering, and SSRF protection across 15 proxied providers. The latency number is measured across 20,000 runs, not claimed.
- Visual rule builder with semantic matching, regex patterns, and PII redaction
- Injection and PII screening on both prompts and responses
- Streaming support for all 15 proxied providers
- Reproducible benchmark: pnpm bench:firewall-latency
Also ships: Shadow AI detection (20 AI providers tracked, credential-leak flags on outbound prompts), AI-SPM posture scoring with 12 misconfiguration checks, and agentless AI-BOM discovery across AWS Bedrock, GCP Vertex AI, and Azure OpenAI.
Read the benchmark04 · Observability + FinOps
Every trace, token, and dollar, accounted per model.
OpenTelemetry-native ingestion turns agent runs into inspectable trajectories, and every request carries its own cost. Budgets, SLA targets, and drift alerts run on the same stream, so spend and quality never drift apart unnoticed.
- OTLP trace ingestion with cost breakdown per trace and per model
- Spend tracking per model, prompt, and team, with budget alerts and forecasting
- SLA targets on availability, latency P95/P99, error rate, and throughput, with automatic violation alerts
- Z-score drift detection triggers re-evaluation automatically, with cooldowns and daily limits against alert fatigue
Also ships: a CISO-level risk dashboard — compliance scores, active incidents with MTTR, and vendor risk on one screen.
See the tracing docs05 · Compliance + Governance
50 frameworks, mapped to evidence instead of screenshots.
Automated assessment across EU AI Act, NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, ISO 42001, India DPDP, HIPAA, and more. Risk classification, gap analysis, and audit-ready documentation generate from your actual runs, not from a questionnaire.
- EU AI Act risk classification with gap analysis and remediation plans
- Structured incident response: RCA templates for 11 AI failure categories, auto severity classification
- Vendor risk: 6-dimension assessments, DPA expiry alerts, lock-in scoring, pre-built profiles for OpenAI, Anthropic, Google, and Mistral
- Regulatory change tracking with alerts, plus PDF export for auditors
06 · Agents + Voice
Agents traced thought by thought. Voice included.
Distributed tracing for multi-step agents: every thought, action, and observation is a span with its own token count and cost. Voice agents run through the same pipeline, with word-level transcription and audio-deepfake detection instead of a separate voice stack.
- Agent trajectory visualization with RAG diagnostics and chunk attribution
- MCP traffic inspection: every tool call risk-scored in real time, memory poisoning flagged
- Autonomy ladder L1–L4 with firewall-guarded deploys
- Voice-specific guardrails and latency evals on the realtime audio pipeline
Also ships: a results copilot that reads scan and eval output, prioritizes findings by risk, and drafts step-by-step fixes with code examples.
Build an agent on the canvasPlatform fundamentals
OpenTelemetry export · BYOK encryption · CI/CD templates · self-hosted Docker · Python SDK · OpenAPI docs
Run it against your own model.
The free tier includes 50,000 traces per month. Or get a walkthrough from the team.