Sourced & dated buyer's guide
AI Gateway & Eval Buyer's Guide.
Side-by-side capability matrix for the 8 most common platforms in AI gateway, eval, and red-team RFPs. Real numbers from each vendor's public docs and source — Portkey and LiteLLM re-derived 2026-08-09, the rest verified 2026-05-07.
SOC 2 evidence engineISO 42001 mappedEU AI ActGDPR
Portkey and LiteLLM cells re-derived from their source on 2026-08-09; the rest were last checked 2026-05-07 and are marked ?where we could not verify them. Vendor comparison pages — including this one — are the weakest kind of source; every cell here is meant to be checkable against the vendor's own repo or docs.
Capability matrix
Eight platforms, one honest scorecard.
| Capability | EvalGuard | Portkey | Langfuse | Helicone | LangSmith | LiteLLM | OpenRouter | Vercel AI Gateway |
|---|---|---|---|---|---|---|---|---|
| Eval scorers (built-in) | 200+ | 0 (evals SDK is OpenAI passthrough) | Evaluator Library (Ragas) + custom | 7 presets + custom LLM-as-judge | Custom + LLM-as-judge | 0 (Evals API proxies OpenAI) | n/a | n/a |
| Red-team plugins | 300+ | 0 (delegated to partners) | 0 | 0 | 0 | 0 | 0 | 0 |
| Attack strategies | 100+ | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Compliance frameworks (in-product mappings, not certifications) | 50 (EU AI Act, ISO 42001, NIST AI RMF, OWASP LLM Top 10, MITRE ATLAS, +45) | 0 framework code mappings (SOC2 / ISO / HIPAA service certs only) | 0 mappings (SOC 2 Type II + ISO 27001 certified) | 0 mappings | 0 mappings (SOC 2 + ISO 27001 certified) | 0 | 0 | 0 |
| Total guardrails (scorers + red-team) | 590 | ~39 (20 deterministic + 19 partner) | Evaluator Library + custom | 7 evaluator presets | Custom | 51 integrations + own content filter | 0 | 0 |
| LLM firewall p95 | 3.67ms (published, /trust/latency) | No published number | n/a | n/a | n/a | n/a | n/a | n/a |
| Provider integrations | 90+ typed | 71 typed ('1,600+ models via aliases') | via LiteLLM (100+) | Proxy-based (broad, per their docs) | via LangChain | 100+ | 200+ | OpenAI + Anthropic + Google |
| Self-hosted (Docker + Helm) | Yes (full platform) | OSS gateway only — observability, prompt mgmt, RBAC, LLM-response cache, OTel exporter, PII model are SaaS-locked | Yes | Yes (per their docs) | Enterprise tier only | Yes | No (SaaS-only) | No (Vercel-tied) |
| OSS license | Apache 2.0 (SDKs + CLI; hosted service proprietary) | Apache 2.0 (gateway frame; control plane closed) | MIT | Apache 2.0 (per their repo) | Closed-source | MIT core; enterprise/ under BerriAI EL | Closed-source | Closed-source |
| Ownership | Independent, bootstrapped | Independent, venture-backed | YC W23, independent | YC W23, independent (per their site) | LangChain Inc subsidiary | Independent | Independent | Vercel platform-tied |
| Cache key normalization | Field-aware by default (role+content+model+tenantId) | SHA256(entire body + URL) — a changed top_p or user forces a miss | n/a | n/a | n/a | Basic | n/a | n/a |
| Conditional/metadata routing DSL | $eq, $ne, $gt, $lt, $in, $nin, $regex, $exists, $contains, $and, $or, $not | Same MongoDB-ish DSL | n/a | n/a | n/a | Basic config rules | Provider preferences | Limited |
| Budget caps with auto-disable | Yes (per-tenant dailyBudgetUsd, enforced inline) | Enterprise tier only | n/a (observability only) | Cost tracking, no auto-disable | n/a | Yes | Per-key spending limit | Vercel platform billing |
| Sticky load balancing | Yes (per-tenant session affinity) | Yes (hosted) | n/a | n/a | n/a | Not found | No | No |
| Quality-cost routing (auto-downgrade for simple queries) | Yes (gateway/quality-cost-router.ts) | No first-class primitive | n/a | n/a | n/a | Cost-aware routing; no quality-based downgrade found | Manual model preferences | No |
| SDK languages | TS, Python, Go, Java | TS, Python | TS, Python | TS, Python | TS, Python, Go, Java | Python | REST only | TS only |
| Public head-to-head benchmarks | Yes (/trust/latency 20K runs, NeMo head-to-head) | None published | None | None | None | None | None | None |
| Pricing transparency | Public ($49/mo Pro, $199/mo Team, /pricing) | Public + Enterprise quote (PII anonymizer + BAA quote-only) | Public | Public | Public + Enterprise quote | Public + Enterprise | Pay-per-token | Vercel platform pricing |
Methodology
- EvalGuard numbers are drift-checked at build time against the live registries (
packages/core/src/counts.ts). Latency is measured at /trust/latency with reproducible scripts. - Portkey numbers come from a code audit of Portkey-AI/gateway at commit 669825cbe (2026-08-09) — including the grep-confirmed absence of red-team and compliance-mapping code, and a recount that moved their provider total from 89 to 71.
- LiteLLM numbers come from the same kind of audit, re-run 2026-08-09. It corrected this page in LiteLLM's favour: they ship 51 guardrail integrations and a real OpenTelemetry exporter, both of which earlier revisions of this table understated.
- Other competitors are from their public docs + GitHubREADMEs. Where marketing was vague (e.g. "1,600+ models"), we counted actual provider directories in their codebases. Helicone has no clone available to us, so its cells rest on their published docs alone.
- Corrections we made against ourselves. A previous revision of this page asserted an acquisition of Portkey by Palo Alto Networks, with a date and a press-release link. No primary source supported it and the claim has been removed. If you find a cell here that is wrong, tell us and we will correct it.
Try EvalGuard Free
No credit card required. Apache 2.0 SDKs + CLI, self-hosted in 5 minutes.