Institutional actuarial evaluation tracking 60+ global AI benchmark suites across 6 domains, B0–B5 lifecycle status, and contamination audits under ZIP-1.0.
Evaluated across Item Response Theory (IRT) Headroom, SOTA Ceilings, Contamination Risk, and Scaffolding Delta.
| Benchmark Suite | Domain | BHS Score | Tier | Ceiling % | Contamination % | IRT Discrim | Harness Sens |
|---|---|---|---|---|---|---|---|
| Humanity's Last Exam (HLE) Center for AI Safety (CAIS) & Scale AI • 2025 | reasoning_math | 98.6 | B5 | 28.5% | 0.8% | 0.96 | 12/100 |
| SWE-bench Pro Princeton NLP & Cognition Labs • 2025 | agentic_coding | 96.2 | B5 | 35.8% | 1.1% | 0.94 | 15.4/100 |
| OSWorld: Multimodal OS Agents Tsinghua University & CMU • 2024 | agentic_general | 94.8 | B5 | 31.4% | 1.5% | 0.92 | 18.2/100 |
| ARC-AGI-2 (Abstraction & Reasoning Corpus 2) François Chollet & ARC Prize Foundation • 2025 | reasoning_math | 95.4 | B5 | 42% | 0.5% | 0.97 | 14.1/100 |
| GDP.pdf: Enterprise PDF Extraction Epoch AI & Independent Actuarial Labs • 2025 | multimodal_vision | 93.7 | B5 | 48.2% | 1.2% | 0.91 | 19.5/100 |
| LiveBench (Monthly Rotating Benchmark) Abacus AI & NYU • 2024 | reasoning_math | 94.1 | B4 | 69.4% | 0.4% | 0.93 | 11.8/100 |
| GSM1k: The Contamination Auditor Scale AI • 2024 | reasoning_math | 93.2 | B4 | 81.5% | 0.2% | 0.89 | 8.5/100 |
| FrontierMath: Research-Grade Mathematics Epoch AI • 2024 | reasoning_math | 95.8 | B4 | 12.5% | 0.1% | 0.98 | 6.2/100 |
| SimpleQA: Factuality Holdout OpenAI • 2024 | knowledge_qa | 91.5 | B4 | 44.2% | 2.1% | 0.88 | 9.4/100 |
| Tau-Bench: Dynamic Customer Service Agents Sierra AI & Stanford • 2024 | agentic_general | 92.4 | B4 | 49.8% | 1.3% | 0.90 | 16.8/100 |
| GPQA Diamond (Google-Proof QA) NYU, Anthropic & Cohere • 2023 | reasoning_math | 89.8 | B3 | 78.4% | 4.2% | 0.94 | 14.5/100 |
| SWE-bench Verified Princeton NLP & OpenAI • 2024 | agentic_coding | 88.5 | B3 | 73.2% | 5.1% | 0.92 | 18.9/100 |
| MATH-500 UC Berkeley & OpenAI • 2024 | reasoning_math | 82.4 | B3 | 94.8% | 8.6% | 0.85 | 12.1/100 |
| AIME 2024 (American Invitational Math Exam) Mathematical Association of America • 2024 | reasoning_math | 91.2 | B3 | 86.7% | 2.4% | 0.95 | 9.8/100 |
| LiveCodeBench UC Berkeley & MIT • 2024 | agentic_coding | 92.8 | B3 | 64.2% | 0.6% | 0.91 | 11.2/100 |
| BigCodeBench-Hard BigCode Project • 2024 | agentic_coding | 87.6 | B3 | 58.4% | 3.8% | 0.89 | 16.4/100 |
| IFEval: Instruction-Following Evaluation Google Research • 2023 | knowledge_qa | 88 | B3 | 89.2% | 4.8% | 0.86 | 8.2/100 |
| MMMU-Pro: Professional Multimodal University of Waterloo & Independent Labs • 2024 | multimodal_vision | 89.4 | B3 | 62.1% | 2.5% | 0.92 | 13.7/100 |
| GAIA: General AI Assistant Meta AI, Hugging Face & AutoGPT • 2023 | agentic_general | 86.8 | B3 | 68.5% | 5.4% | 0.89 | 22.4/100 |
| MathVista: Visual Mathematical Reasoning UCLA & University of Washington • 2023 | multimodal_vision | 83.5 | B3 | 76.8% | 7.2% | 0.86 | 15.2/100 |
| SWE-bench Lite Princeton NLP • 2024 | agentic_coding | 62.4 | B2 | 54% | 12.8% | 0.76 | 28.4/100 |
| LMSYS Chatbot Arena (Elo) LMSYS Org (UC Berkeley) • 2023 | knowledge_qa | 64.8 | B2 | 91% | 8.5% | 0.78 | 26.5/100 |
| AlpacaEval 2.0 Stanford CRFM • 2024 | knowledge_qa | 61 | B2 | 88.5% | 9.1% | 0.74 | 25.8/100 |
| MMLU-Redux Edgar Chen et al. • 2024 | knowledge_qa | 68.5 | B2 | 91.5% | 11.2% | 0.79 | 16.5/100 |
| EvalPlus (HumanEval+) EvalPlus Team • 2023 | agentic_coding | 65.2 | B2 | 92.4% | 16.5% | 0.75 | 14.8/100 |
| MMLU (Massive Multitask Language Understanding) UC Berkeley (Dan Hendrycks et al.) • 2020 | knowledge_qa | 41.2 | B1 | 92.3% | 28.5% | 0.52 | 19.8/100 |
| GSM8K (Grade School Math 8K) OpenAI (Cobbe et al.) • 2021 | reasoning_math | 38.5 | B1 | 97.8% | 34.2% | 0.44 | 12.4/100 |
| HellaSwag: Common Sense NLI Rowan Zellers et al. (UW / AI2) • 2019 | knowledge_qa | 32 | B1 | 98.4% | 42% | 0.38 | 11.5/100 |
| ARC-Challenge (AI2 Reasoning Challenge) Allen Institute for AI (AI2) • 2018 | reasoning_math | 36.4 | B1 | 96.8% | 38.6% | 0.41 | 10.2/100 |
| TruthfulQA Lin et al. (Oxford / AI Safety) • 2021 | safety_alignment | 42.6 | B1 | 88% | 31% | 0.48 | 24.1/100 |
Licensed under CC BY 4.0. Permitted for academic, enterprise, and search engine citation.