Benchmarks
28 of EvalGuard's 34 benchmark suites detailed here — 17 academic knowledge + reasoning suites and 11 safety/adversarial suites. Each ships a runnable case bank, a deterministic scorer, and reproducible runs.
--from-hf; otherwise only the authored bank exists. Full per-dataset licensing is in DATASET-MODEL-LICENSES.md.Why benchmark coverage matters
A named suite gives a regression a shape: “faithfulness dropped” is an argument, “MMLU-format knowledge accuracy fell 6 points between these two prompt versions” is a number two engineers can act on. EvalGuard covers the seven task formats DeepEval markets (MMLU, HellaSwag, BIG-Bench Hard, DROP, TruthfulQA, HumanEval, GSM8K) plus 16 more academic suites — including medical, legal, finance, security and multimodal — and 11 safety suites. What that buys is comparability across your own runs: same questions, same scorer, same seed, every time. It does not buy comparability with a published leaderboard, because the cases are ours, not the public datasets'. Run any suite with one command:
evalguard benchmark run mmlu --model gpt-4o evalguard benchmark run truthfulqa --model claude-opus-4 evalguard benchmark run humaneval --model gemini-2.5-pro
Academic & reasoning benchmarks (17)
Knowledge, reasoning, code, math, and vertical-domain suites for comparing models against each other and against your own earlier runs. Each reproduces the task format of the public benchmark it is named after; the cases are EvalGuard-written, and the count on each card is the count that suite ships.
MMLU
General knowledge
Massive Multitask Language Understanding — 57 academic subjects from elementary school through professional-level (US history, computer science, law, medicine, ethics, etc.).
BIG-Bench Hard (BBH)
Reasoning
Multi-step reasoning tasks in the BIG-Bench Hard style — the family where the prior-best LM averaged below the average human rater.
DROP
Reading comprehension
Discrete Reasoning Over Paragraphs — requires multi-step arithmetic, counting, sorting over text-extracted facts.
BoolQ
Reading comprehension
Yes/no question-passage pairs in the BoolQ style, drawn from natural search-query phrasing.
TruthfulQA
Factuality
Questions engineered to elicit imitative falsehoods — urban legends, common misconceptions, conspiracy theories.
HellaSwag
Commonsense reasoning
Multiple-choice sentence completion testing commonsense reasoning. Designed adversarially so humans score 95%+ but models often fail.
HumanEval
Code
Hand-written Python programming problems in the HumanEval style; pass@1 measured by running unit tests against generated code.
GSM8K
Math
Grade-school math word problems requiring 2-8 step reasoning chains.
ARC
Science reasoning
Grade-school science questions in the ARC style, spanning the Easy and Challenge difficulty split.
BBQ
Bias
Bias Benchmark for QA — measures social biases in 9 demographic categories (age, disability, gender identity, nationality, physical appearance, race/ethnicity, religion, SES, sexual orientation).
IFEval
Instruction following
Verifiable instruction-following — measures whether a model follows specific formatting / length / structure constraints in a prompt.
MMMU
Multimodal
Multi-discipline questions requiring combined image + text reasoning, in the MMMU style.
VisionBench
Multimodal
Curated suite covering image-question reasoning, OCR, chart interpretation, scientific figure comprehension.
MedQA
Vertical — Medical
USMLE-style medical questions covering disease diagnosis, treatment and clinical ethics.
LegalBench
Vertical — Legal
Legal-reasoning tasks in the LegalBench style: issue spotting, rule recall, application, and conclusions.
Financial Reasoning Benchmark
Vertical — Finance
Financial reasoning over calculations, accounting, market analysis, risk, and compliance. EvalGuard-authored synthetic corpus; the external FinanceBench dataset is not bundled.
CyberBench
Vertical — Security
Cybersecurity question bank across pentesting, vulnerability classification, threat intelligence, CWE/CVE matching.
Safety & adversarial benchmarks (11)
Adversarial datasets for measuring refusal quality, jailbreak resistance, and harm-category coverage. Each is referenced by ID in our red-team plugin registry (see /docs/plugins) so they double as both standalone benchmarks AND first-class red-team plugins.
AEGIS
Content safety
Content-safety probes across 13 risk categories. EvalGuard-authored synthetic corpus, following the AEGIS risk taxonomy — the dataset itself is not bundled.
Harmful-Content Refusal
Harmful content
Refusal probes across 14 harm categories. EvalGuard-authored synthetic corpus, inspired by the BeaverTails taxonomy — the dataset itself is not bundled.
HarmBench
Adversarial
Standardized harmful-behavior probes across 7 categories. EvalGuard-authored synthetic corpus, following the HarmBench behaviour taxonomy — the dataset itself is not bundled.
Pliny / L1B3RT4S
Jailbreak
Jailbreak-structure attacks — persona shift, chain-of-thought hijack, system-prompt extraction, reflective jailbreak. EvalGuard-authored re-implementations of the four canonical Pliny families; the L1B3RT4S archive is deliberately not bundled.
Toxic-Conversation Robustness
Production toxicity
Toxic-conversation robustness probes modeling production-distribution attack patterns. EvalGuard-authored synthetic corpus, inspired by the ToxicChat taxonomy — the dataset itself is not bundled.
CyberSecEval
Code security
Secure-code-generation and cyber-assistance-refusal probes spanning common CWEs and MITRE ATT&CK categories. EvalGuard-authored synthetic corpus, following the CyberSecEval taxonomy — the dataset itself is not bundled.
Unsafe-Content Safety
Multimodal safety
Unsafe-content safety probes across 11 categories. EvalGuard-authored synthetic corpus, inspired by the UnsafeBench taxonomy — the dataset itself is not bundled.
VLGuard
Multimodal safety
Text re-grounding of the VLGuard safety taxonomy — privacy, deception, discrimination, dangerous actions — so a text-only firewall is exercised on the same harm categories. EvalGuard-authored.
VLSU
Multimodal safety
Text re-grounding of the VLSU underrepresented-harm categories: cultural erasure, minority-targeted misinformation, low-resource-language jailbreaks, disability stereotyping. EvalGuard-authored.
Do-Not-Answer
Refusal calibration
Prompts an LLM should decline or carefully caveat, across the Do-Not-Answer risk areas. EvalGuard-authored synthetic corpus — the LibrAI dataset itself is not bundled.
XSTest
Over-refusal
Safe prompts that LOOK harmful, plus unsafe contrasts. EvalGuard-authored synthetic corpus, following the XSTest exaggerated-safety design — the dataset itself is not bundled.
Public safety corpus — via the firewall-benchmark API
Apache-2.0 · public sample split · 10 source taxonomies
A unified, license-clean slice of the safety corpus is served from the EvalGuard firewall-benchmark API so the research community can benchmark refusal calibration, jailbreak resistance, and content-policy behaviour against the same prompts EvalGuard uses internally. The public sample split is fully visible; the held-out split is served as metadata only (category list + counts).
GET /api/v1/firewall-benchmark/public-corpusEvery prompt in the public corpus is original EvalGuard staff writing, inspired by 10 published safety-research taxonomies (AEGIS, BeaverTails, CyberSecEval, DoNotAnswer, HarmBench, ToxicChat, UnsafeBench, VLGuard, VLSU, XSTest). We license-audited the source corpora before release — the Pliny / L1B3RT4S jailbreak archive is intentionally not included.
Fetch it over HTTP (no API key required):
curl https://evalguard.ai/api/v1/firewall-benchmark/public-corpus
# Response: { meta, publicCorpus: [{ input, category, source, expectedBehavior }], note }
# For each prompt, check whether your model's response matches expectedBehavior
# ("block" for prompts that should be refused, "allow" for safe prompts).CLI + SDK
Every benchmark is runnable from the CLI:
# Run a single benchmark (suite is a positional arg) evalguard benchmark run mmlu --model gpt-4o # Compare models — run the same benchmark once per model evalguard benchmark run humaneval --model gpt-4o evalguard benchmark run humaneval --model claude-opus-4 # Save the report (emit JSON with --json and redirect stdout) evalguard benchmark run gsm8k --model gpt-4o --json > report.json
The CLI runs the suite locally (via @evalguard/core), so it needs no API key. The cross-model leaderboard — aggregated from eval and security-scan data — is also available from the Python SDK:
from evalguard import EvalGuardClient
client = EvalGuardClient(api_key="eg_live_...")
board = client.get_leaderboard(category="overall")
print(board["leaderboard"]) # [{"model": "gpt-4o", "overallScore": 0.87, ...}, ...]