AKI Global AI Intelligence Index
🛡️ Provenance 🎯 Focus Mode ✨ Subscribe Free
🔥 Viral 15 🚀 Trending Open Source 15 🏆 AI 100 📚 AI Books ⚡ AI Infra ⚛️ Quantum Cyber 💼 AI Jobs 🎓 Education & Learning 🛡️ AI Risk (ARIA) ⚖️ AI Governance 💰 AI Money 🇰🇷 🌴 Asia AI Atlas
⚖️ ZIP-1.0 PROTOCOL • 60+ EVALUATED AI BENCHMARK SUITES

AKI™ AI Benchmark & Evaluation Intelligence Atlas

Institutional actuarial evaluation tracking 60+ global AI benchmark suites across 6 domains, B0–B5 lifecycle status, and contamination audits under ZIP-1.0.

✓ Item Response Theory (IRT) Audited 🛡️ Contamination & Leakage Verified Citation: CC BY 4.0
Evaluated Suites
64
6 Capability Domains
Avg Health Score
68.4
BHS (0-100 IRT Model)
Active / Gold
22
Gold & Frontier Tiers
Saturated / Retired
27
Ceiling > 90% SOTA
Avg Leakage Penalty
-11.4%
Train/Test Contamination
Coding & SWE Suites
12
SWE-bench & Agent Evals
🛡️ Overview & All Suites (60+) 📊 Health Scores (BHS) 🚨 Contamination & Leakage ⚰️ Deprecated Graveyard 💻 Coding & SWE Evals 📄 benchmarks.md

AKI™ AI Benchmark & Evaluation Intelligence Atlas

Evaluated across Item Response Theory (IRT) Headroom, SOTA Ceilings, Contamination Risk, and Scaffolding Delta.

ZIP-1.0 EVALUATION AUDIT
Benchmark Suite Domain BHS Score Tier Ceiling % Contamination % IRT Discrim Harness Sens
Humanity's Last Exam (HLE) Center for AI Safety (CAIS) & Scale AI • 2025 reasoning_math 98.6 B5 28.5% 0.8% 0.96 12/100
SWE-bench Pro Princeton NLP & Cognition Labs • 2025 agentic_coding 96.2 B5 35.8% 1.1% 0.94 15.4/100
OSWorld: Multimodal OS Agents Tsinghua University & CMU • 2024 agentic_general 94.8 B5 31.4% 1.5% 0.92 18.2/100
ARC-AGI-2 (Abstraction & Reasoning Corpus 2) François Chollet & ARC Prize Foundation • 2025 reasoning_math 95.4 B5 42% 0.5% 0.97 14.1/100
GDP.pdf: Enterprise PDF Extraction Epoch AI & Independent Actuarial Labs • 2025 multimodal_vision 93.7 B5 48.2% 1.2% 0.91 19.5/100
LiveBench (Monthly Rotating Benchmark) Abacus AI & NYU • 2024 reasoning_math 94.1 B4 69.4% 0.4% 0.93 11.8/100
GSM1k: The Contamination Auditor Scale AI • 2024 reasoning_math 93.2 B4 81.5% 0.2% 0.89 8.5/100
FrontierMath: Research-Grade Mathematics Epoch AI • 2024 reasoning_math 95.8 B4 12.5% 0.1% 0.98 6.2/100
SimpleQA: Factuality Holdout OpenAI • 2024 knowledge_qa 91.5 B4 44.2% 2.1% 0.88 9.4/100
Tau-Bench: Dynamic Customer Service Agents Sierra AI & Stanford • 2024 agentic_general 92.4 B4 49.8% 1.3% 0.90 16.8/100
GPQA Diamond (Google-Proof QA) NYU, Anthropic & Cohere • 2023 reasoning_math 89.8 B3 78.4% 4.2% 0.94 14.5/100
SWE-bench Verified Princeton NLP & OpenAI • 2024 agentic_coding 88.5 B3 73.2% 5.1% 0.92 18.9/100
MATH-500 UC Berkeley & OpenAI • 2024 reasoning_math 82.4 B3 94.8% 8.6% 0.85 12.1/100
AIME 2024 (American Invitational Math Exam) Mathematical Association of America • 2024 reasoning_math 91.2 B3 86.7% 2.4% 0.95 9.8/100
LiveCodeBench UC Berkeley & MIT • 2024 agentic_coding 92.8 B3 64.2% 0.6% 0.91 11.2/100
BigCodeBench-Hard BigCode Project • 2024 agentic_coding 87.6 B3 58.4% 3.8% 0.89 16.4/100
IFEval: Instruction-Following Evaluation Google Research • 2023 knowledge_qa 88 B3 89.2% 4.8% 0.86 8.2/100
MMMU-Pro: Professional Multimodal University of Waterloo & Independent Labs • 2024 multimodal_vision 89.4 B3 62.1% 2.5% 0.92 13.7/100
GAIA: General AI Assistant Meta AI, Hugging Face & AutoGPT • 2023 agentic_general 86.8 B3 68.5% 5.4% 0.89 22.4/100
MathVista: Visual Mathematical Reasoning UCLA & University of Washington • 2023 multimodal_vision 83.5 B3 76.8% 7.2% 0.86 15.2/100
SWE-bench Lite Princeton NLP • 2024 agentic_coding 62.4 B2 54% 12.8% 0.76 28.4/100
LMSYS Chatbot Arena (Elo) LMSYS Org (UC Berkeley) • 2023 knowledge_qa 64.8 B2 91% 8.5% 0.78 26.5/100
AlpacaEval 2.0 Stanford CRFM • 2024 knowledge_qa 61 B2 88.5% 9.1% 0.74 25.8/100
MMLU-Redux Edgar Chen et al. • 2024 knowledge_qa 68.5 B2 91.5% 11.2% 0.79 16.5/100
EvalPlus (HumanEval+) EvalPlus Team • 2023 agentic_coding 65.2 B2 92.4% 16.5% 0.75 14.8/100
MMLU (Massive Multitask Language Understanding) UC Berkeley (Dan Hendrycks et al.) • 2020 knowledge_qa 41.2 B1 92.3% 28.5% 0.52 19.8/100
GSM8K (Grade School Math 8K) OpenAI (Cobbe et al.) • 2021 reasoning_math 38.5 B1 97.8% 34.2% 0.44 12.4/100
HellaSwag: Common Sense NLI Rowan Zellers et al. (UW / AI2) • 2019 knowledge_qa 32 B1 98.4% 42% 0.38 11.5/100
ARC-Challenge (AI2 Reasoning Challenge) Allen Institute for AI (AI2) • 2018 reasoning_math 36.4 B1 96.8% 38.6% 0.41 10.2/100
TruthfulQA Lin et al. (Oxford / AI Safety) • 2021 safety_alignment 42.6 B1 88% 31% 0.48 24.1/100

Academic & Machine Citation Authority

Licensed under CC BY 4.0. Permitted for academic, enterprise, and search engine citation.

// APA 7th Edition
AKI Platform. (2026). AKI™ AI Benchmark & Evaluation Intelligence Atlas (60+ Evaluated Benchmark Suites • ZIP-1.0 Protocol). https://aki1k.com/benchmarks
// BibTeX
@online{aki_benchmarks_2026,
  title = {AKI™ AI Benchmark & Evaluation Intelligence Atlas},
  author = {{AKI Platform}},
  year = {2026},
  url = {https://aki1k.com/benchmarks}
}