Top 20 LLM evaluation tools · 2026.
20 LLM evaluation, security, and observability tools, ranked. Every entry carries its own verification date, none older than 2026-07-22 — /compare carries dated feature matrices for these and for the gateway, tracing and AI-security tools not ranked here.
Evaluation Frameworks
EvalGuard
RecommendedThe all-in-one AI evaluation and security platform
300+ attack plugins, 200+ scorers, 77 LLM providers, compliance dashboard, LLM firewall, and full SaaS platform. SDKs and CLI are Apache 2.0.
Promptfoo
Open-source LLM eval, acquired by OpenAI (announced 2026-03-09)
Verified 2026-08-09
Popular open-source evaluation framework with 155 red team plugins, 66 assertion types, and 60+ providers. OpenAI acquisition announced 2026-03-09; they state the project stays open source and MIT. ~24K GitHub stars, 300K+ developers. Free Community tier; Enterprise and On-Premise pricing is not published.
DeepEval / Confident AI
Python-native eval framework with growing red team
Verified 2026-08-09
Python-first eval framework with 50 metrics and 50+ vulnerability types (via DeepTeam). Native pytest integration. ~17.5K GitHub stars, Apache-2.0, 400K+ monthly downloads. Free OSS, Confident AI from $200/mo.
Braintrust
Closed control plane, MIT proxy and scorer library
Verified 2026-08-10
AI evaluation platform focused on production eval workflows and CI/CD integration. Control plane is closed, but the AI Proxy and autoevals scorer library are MIT and the data plane can be self-hosted in your own AWS account. Free tier; Pro $249/mo; Enterprise custom.
MLflow
Databricks' ML lifecycle platform
Verified 2026-08-10
Open-source ML lifecycle management with basic LLM eval. SaaS requires Databricks. No security testing.
Security & Red Teaming
Giskard
EU-focused red teaming with adaptive agents
Verified 2026-08-10
European open-source AI red teaming platform. The 3.0 rewrite registers 7 vulnerability generators and ~19 checks including 8 LLM judges, with dynamic multi-turn red teaming and Logfire tracing. 1 compliance framework (OWASP LLM Top 10).
Garak (NVIDIA)
NVIDIA's LLM vulnerability scanner
Verified 2026-08-09
Open-source LLM vulnerability scanner with 190 probes across 43 modules and 117 detectors. CLI only, no SaaS platform.
PyRIT (Microsoft)
Microsoft's red team dev library
Verified 2026-08-09
Python Risk Identification Toolkit for generative AI. Developer library, not a platform.
Mindgard
Enterprise AI security for SOC teams
Verified 2026-08-12
Enterprise AI security platform with MITRE ATLAS alignment, built on Lancaster University AI security research. Targets enterprise SOC teams; deploys via CI/CD, Burp Suite and APIs.
Lakera (Check Point)
Enterprise LLM firewall (sub-50ms claimed), plus an AI red teaming product
Verified 2026-08-12
AI security platform with an enterprise LLM firewall (sub-50ms latency claimed by Lakera; EvalGuard publishes 3.67ms p95 measured at /trust/latency), proprietary threat intelligence, and a dedicated AI Red Teaming product for continuous adversarial testing. Acquired by Check Point. Closed platform — capabilities it does not advertise cannot be verified. Free (10K req/mo), Enterprise custom.
Purple Llama (Meta)
Meta's safety benchmarks and Llama Guard
Verified 2026-07-22
Meta's open-source AI safety initiative with CyberSecEval and Llama Guard. Benchmarks and models, not a platform.
Observability & Monitoring
Langfuse
Best-in-class open-source LLM observability (YC W23)
Verified 2026-08-09
Leading open-source LLM observability platform with best-in-class tracing, prompt management, and 100+ providers via LiteLLM. It ships an Evaluator Library (LLM-as-a-judge, built with Ragas) but no red teaming and no runtime firewall. Hobby tier free (50K units/mo, 30-day access); Core $29/mo; Pro $199/mo.
Maxim AI
End-to-end AI evaluation and observability
Verified 2026-08-10
End-to-end AI evaluation and observability with agent simulation, tracing, cost tracking. Publishes SOC 2 and ISO 27001 audit status plus HIPAA and GDPR programmes. Free Developer tier (3 seats); Professional $29/seat/mo; Business $49/seat/mo.
Arize AI / Phoenix
Free-tier observability (Phoenix OSS is source-available, ELv2)
Verified 2026-08-10
Arize sells the managed AX platform alongside Phoenix, a source-available (Elastic License 2.0) observability tool with pre-built evaluators, a prompt playground and versioned prompt management. 10.7K GitHub stars, 2.5M+ downloads. No red teaming, no attack plugin library, no firewall, no gateway.
Datadog LLM Observability
Infrastructure monitoring giant adds LLM features
Verified 2026-08-10
Industry-leading monitoring platform that now ships LLM observability with built-in evaluations, plus AI Guard, an inline runtime guardrail against prompt injection, jailbreaking and tool misuse. Counts for both are unpublished.
Weights & Biases
ML experiment tracking with Weave for LLMs
Verified 2026-08-10
Leading ML experiment tracking platform. Weave adds basic LLM evaluation but no security testing.
Big Tech (Vendor-Locked)
OpenAI Evals
Free eval, locked to OpenAI models
Verified 2026-07-22
OpenAI's built-in evaluation framework. Free for OpenAI users but completely locked to the OpenAI ecosystem.
Google Vertex AI Evaluation
GCP-only evaluation tools
Verified 2026-07-22
Built-in model evaluation on Google Cloud. Works with Google models only, no standalone usage.
Azure AI Content Safety
Content filtering locked to Azure
Verified 2026-07-22
Azure's content moderation and prompt shielding. Strong content filtering but limited to Azure ecosystem.
Consulting Tools
ARTKIT (BCG)
BCG's red teaming Python library
Verified 2026-08-10
BCG X's open-source toolkit for automated red teaming. Python library only, no SaaS or enterprise features.
The verdict
Why teams choose EvalGuard
300+ attack plugins. 200+ eval scorers. 77 LLM providers. Compliance dashboard. LLM firewall. All in one open-source platform.
Competitor data (GitHub stars, downloads, feature counts, pricing, acquisition status) is checked against each vendor’s own live pricing or docs page, or against their published source, and every entry above states the date it was last checked — none older than 2026-07-22. Figures we could not verify against a first-party source have been removed rather than carried forward. If something here is out of date, tell us and we will correct it.