Skip to content
2026 Guide

Top 20 LLM evaluation tools · 2026. 

20 LLM evaluation, security, and observability tools, ranked. Every entry carries its own verification date, none older than 2026-07-22 — /compare carries dated feature matrices for these and for the gateway, tracing and AI-security tools not ranked here.

SOC 2 evidence engineISO 42001 mappedEU AI ActGDPR

Evaluation Frameworks

EvalGuard

Recommended

The all-in-one AI evaluation and security platform

300+ attack plugins, 200+ scorers, 77 LLM providers, compliance dashboard, LLM firewall, and full SaaS platform. SDKs and CLI are Apache 2.0.

Widest attack + eval coverage in this guideFull SaaS + self-hostedEU AI Act + ISO 42001 control mappingsYounger project than the OSS incumbentsNo model registryApache 2.0 covers the SDKs + CLI, not the hosted service

Promptfoo

Open-source LLM eval, acquired by OpenAI (announced 2026-03-09)

Verified 2026-08-09

Popular open-source evaluation framework with 155 red team plugins, 66 assertion types, and 60+ providers. OpenAI acquisition announced 2026-03-09; they state the project stays open source and MIT. ~24K GitHub stars, 300K+ developers. Free Community tier; Enterprise and On-Premise pricing is not published.

Large community (300K+ devs)155 attack pluginsGood CI/CD templatesStill MIT under OpenAIOwned by an LLM vendorFree tier caps red teaming at 10K probes/monthEnterprise + On-Premise pricing not published

DeepEval / Confident AI

Python-native eval framework with growing red team

Verified 2026-08-09

Python-first eval framework with 50 metrics and 50+ vulnerability types (via DeepTeam). Native pytest integration. ~17.5K GitHub stars, Apache-2.0, 400K+ monthly downloads. Free OSS, Confident AI from $200/mo.

Native pytest integration~17.5K GitHub stars50 metricsOTel tracing + 7 DeepTeam guardrailsPython-first (TS SDK only via Confident AI)50+ vulnerability types (vs 300+)No LLM gateway

Braintrust

Closed control plane, MIT proxy and scorer library

Verified 2026-08-10

AI evaluation platform focused on production eval workflows and CI/CD integration. Control plane is closed, but the AI Proxy and autoevals scorer library are MIT and the data plane can be self-hosted in your own AWS account. Free tier; Pro $249/mo; Enterprise custom.

Polished eval UXMIT AI Proxy + autoevalsCI/CD integrationControl plane closed; AI Proxy + autoevals MITNo attack pluginsSelf-hosting is hybrid (data plane only)

MLflow

Databricks' ML lifecycle platform

Verified 2026-08-10

Open-source ML lifecycle management with basic LLM eval. SaaS requires Databricks. No security testing.

Mature model registryDatabricks ecosystemLarge OSS community24 built-in GenAI scorers + judgesNo attack pluginsSaaS needs Databricks

Security & Red Teaming

Giskard

EU-focused red teaming with adaptive agents

Verified 2026-08-10

European open-source AI red teaming platform. The 3.0 rewrite registers 7 vulnerability generators and ~19 checks including 8 LLM judges, with dynamic multi-turn red teaming and Logfire tracing. 1 compliance framework (OWASP LLM Top 10).

Dynamic multi-turn red teamingCompliance-focused toolingEstablished enterprise adoption7 vulnerability generators (vs 300+)~19 checksNo firewall or gateway

Garak (NVIDIA)

NVIDIA's LLM vulnerability scanner

Verified 2026-08-09

Open-source LLM vulnerability scanner with 190 probes across 43 modules and 117 detectors. CLI only, no SaaS platform.

Backed by NVIDIAOpen source190 probes, 117 detectorsCLI onlyNo general-purpose scorer libraryNo dashboard

PyRIT (Microsoft)

Microsoft's red team dev library

Verified 2026-08-09

Python Risk Identification Toolkit for generative AI. Developer library, not a platform.

Backed by Microsoft14 attack strategies + 90 convertersScoring package includedDev library onlyNo dashboardNo built-in scorer catalogue

Mindgard

Enterprise AI security for SOC teams

Verified 2026-08-12

Enterprise AI security platform with MITRE ATLAS alignment, built on Lancaster University AI security research. Targets enterprise SOC teams; deploys via CI/CD, Burp Suite and APIs.

Enterprise SOC focusMITRE ATLAS alignmentCI/CD + Burp Suite integrationsProduct source not publishedPricing tiers not publishedClosed platform — capabilities cannot be verified from source

Lakera (Check Point)

Enterprise LLM firewall (sub-50ms claimed), plus an AI red teaming product

Verified 2026-08-12

AI security platform with an enterprise LLM firewall (sub-50ms latency claimed by Lakera; EvalGuard publishes 3.67ms p95 measured at /trust/latency), proprietary threat intelligence, and a dedicated AI Red Teaming product for continuous adversarial testing. Acquired by Check Point. Closed platform — capabilities it does not advertise cannot be verified. Free (10K req/mo), Enterprise custom.

LLM firewall (sub-50ms claimed)AI Red Teaming productProprietary threat intelCheck Point backingClosed — no published counts to compareFirewall latency claimed, not published as measuredTracing, prompt IDE and compliance coverage not published

Purple Llama (Meta)

Meta's safety benchmarks and Llama Guard

Verified 2026-07-22

Meta's open-source AI safety initiative with CyberSecEval and Llama Guard. Benchmarks and models, not a platform.

Backed by MetaLlama Guard modelCyberSecEvalNot a platformLlama-focusedNo dashboard

Observability & Monitoring

Langfuse

Best-in-class open-source LLM observability (YC W23)

Verified 2026-08-09

Leading open-source LLM observability platform with best-in-class tracing, prompt management, and 100+ providers via LiteLLM. It ships an Evaluator Library (LLM-as-a-judge, built with Ragas) but no red teaming and no runtime firewall. Hobby tier free (50K units/mo, 30-day access); Core $29/mo; Pro $199/mo.

Best-in-class tracing100+ providers (LiteLLM)Good prompt managementSOC 2 Type II + ISO 27001 certifiedZero attack pluginsEvaluator library is smaller and partner-ledNo red teaming or firewall

Maxim AI

End-to-end AI evaluation and observability

Verified 2026-08-10

End-to-end AI evaluation and observability with agent simulation, tracing, cost tracking. Publishes SOC 2 and ISO 27001 audit status plus HIPAA and GDPR programmes. Free Developer tier (3 seats); Professional $29/seat/mo; Business $49/seat/mo.

Agent simulationSOC 2 + ISO 27001 auditedCost trackingPlatform closed; Bifrost gateway Apache-2.0No first-party red-team plugin libraryGuardrails are mostly third-party integrations

Arize AI / Phoenix

Free-tier observability (Phoenix OSS is source-available, ELv2)

Verified 2026-08-10

Arize sells the managed AX platform alongside Phoenix, a source-available (Elastic License 2.0) observability tool with pre-built evaluators, a prompt playground and versioned prompt management. 10.7K GitHub stars, 2.5M+ downloads. No red teaming, no attack plugin library, no firewall, no gateway.

Free AX tier (25K spans/mo, 15-day retention)10.7K stars, 2.5M+ downloadsPrompt playground + pre-built evaluatorsZero attack pluginsNo firewall/gatewayPhoenix is ELv2, not OSI open source

Datadog LLM Observability

Infrastructure monitoring giant adds LLM features

Verified 2026-08-10

Industry-leading monitoring platform that now ships LLM observability with built-in evaluations, plus AI Guard, an inline runtime guardrail against prompt injection, jailbreaking and tool misuse. Counts for both are unpublished.

Best-in-class monitoringVery large enterprise install baseDeep APMBuilt-in evals + AI GuardNo documented attack-plugin libraryEval scorer count not published$35+/host/month

Weights & Biases

ML experiment tracking with Weave for LLMs

Verified 2026-08-10

Leading ML experiment tracking platform. Weave adds basic LLM evaluation but no security testing.

Best experiment trackingModel registryLarge communityNo red-team plugins24 eval scorers (Weave)No compliance mappings

Big Tech (Vendor-Locked)

OpenAI Evals

Free eval, locked to OpenAI models

Verified 2026-07-22

OpenAI's built-in evaluation framework. Free for OpenAI users but completely locked to the OpenAI ecosystem.

Free for OpenAI usersDeep GPT integrationOpenAI models onlyNo red teamingVendor locked

Google Vertex AI Evaluation

GCP-only evaluation tools

Verified 2026-07-22

Built-in model evaluation on Google Cloud. Works with Google models only, no standalone usage.

Free on GCPGemini integrationAutoMLGCP onlyNo red teamingVendor locked

Azure AI Content Safety

Content filtering locked to Azure

Verified 2026-07-22

Azure's content moderation and prompt shielding. Strong content filtering but limited to Azure ecosystem.

Enterprise content filteringAzure compliancePrompt shieldsAzure ecosystem onlyContent moderation + prompt shields, not an evaluation suiteNo red-team plugin library

Consulting Tools

ARTKIT (BCG)

BCG's red teaming Python library

Verified 2026-08-10

BCG X's open-source toolkit for automated red teaming. Python library only, no SaaS or enterprise features.

BCG backingStructured testingOpen sourceNo built-in attack library (pipeline framework)Python library onlyNo dashboard

The verdict

Why teams choose EvalGuard

300+ attack plugins. 200+ eval scorers. 77 LLM providers. Compliance dashboard. LLM firewall. All in one open-source platform.

Competitor data (GitHub stars, downloads, feature counts, pricing, acquisition status) is checked against each vendor’s own live pricing or docs page, or against their published source, and every entry above states the date it was last checked — none older than 2026-07-22. Figures we could not verify against a first-party source have been removed rather than carried forward. If something here is out of date, tell us and we will correct it.