AKI Global AI Intelligence Index
🛡️ Provenance 🎯 Focus Mode ✨ Subscribe Free
🔥 Viral 15 🚀 Trending Open Source 15 🏆 AI 100 📚 AI Books ⚡ AI Infra ⚛️ Quantum Cyber 💼 AI Jobs 🎓 Education & Learning 🛡️ AI Risk (ARIA) ⚖️ AI Governance 💰 AI Money 🇰🇷 🌴 Asia AI Atlas
⚖️ ZIP-1.0 PROTOCOL • 60+ EVALUATED AI BENCHMARK SUITES

AKI™ AI Benchmark & Evaluation Intelligence Atlas

Full audit of SWE-bench Verified, SWE-bench Pro, SWE-bench Multimodal, RepoQA, and LiveCodeBench measuring real developer agent autonomy and test pass rates.

✓ Item Response Theory (IRT) Audited 🛡️ Contamination & Leakage Verified Citation: CC BY 4.0
Evaluated Suites
64
6 Capability Domains
Avg Health Score
68.4
BHS (0-100 IRT Model)
Active / Gold
22
Gold & Frontier Tiers
Saturated / Retired
27
Ceiling > 90% SOTA
Avg Leakage Penalty
-11.4%
Train/Test Contamination
Coding & SWE Suites
12
SWE-bench & Agent Evals
🛡️ Overview & All Suites (60+) 📊 Health Scores (BHS) 🚨 Contamination & Leakage ⚰️ Deprecated Graveyard 💻 Coding & SWE Evals 📄 benchmarks.md

Agentic Coding & Software Engineering Evaluation Matrix

Evaluated across Item Response Theory (IRT) Headroom, SOTA Ceilings, Contamination Risk, and Scaffolding Delta.

ZIP-1.0 EVALUATION AUDIT
Benchmark Suite Domain BHS Score Tier Ceiling % Contamination % IRT Discrim Harness Sens
SWE-bench Pro Princeton NLP & Cognition Labs • 2025 agentic_coding 96.2 B5 35.8% 1.1% 0.94 15.4/100
SWE-bench Verified Princeton NLP & OpenAI • 2024 agentic_coding 88.5 B3 73.2% 5.1% 0.92 18.9/100
LiveCodeBench UC Berkeley & MIT • 2024 agentic_coding 92.8 B3 64.2% 0.6% 0.91 11.2/100
BigCodeBench-Hard BigCode Project • 2024 agentic_coding 87.6 B3 58.4% 3.8% 0.89 16.4/100
SWE-bench Lite Princeton NLP • 2024 agentic_coding 62.4 B2 54% 12.8% 0.76 28.4/100
EvalPlus (HumanEval+) EvalPlus Team • 2023 agentic_coding 65.2 B2 92.4% 16.5% 0.75 14.8/100
HumanEval (Original 164 Problems) OpenAI (Chen et al.) • 2021 agentic_coding 24.5 B0 96.5% 55% 0.28 14.2/100
MBPP (Mostly Basic Python Problems) Google Research (Austin et al.) • 2021 agentic_coding 28.6 B0 95.2% 48.5% 0.32 11/100
Cybench UC Berkeley & NYU • 2024 agentic_coding 94.2 B5 22.8% 0.5% 0.89 2.3/100
InterCode Princeton NLP • 2023 agentic_coding 84.5 B3 72% 4.5% 0.80 6.2/100
RepoBench HKUST • 2023 agentic_coding 74 B2 84.5% 9.8% 0.70 10.4/100
DeepBench Baidu Research • 2017 agentic_coding 11 B0 99.5% 80% 0.10 35.6/100

Academic & Machine Citation Authority

Licensed under CC BY 4.0. Permitted for academic, enterprise, and search engine citation.

// APA 7th Edition
AKI Platform. (2026). Agentic Coding & Software Engineering Evaluation Matrix (60+ Evaluated Benchmark Suites • ZIP-1.0 Protocol). https://aki1k.com/benchmarks/coding
// BibTeX
@online{aki_benchmarks_2026,
  title = {Agentic Coding & Software Engineering Evaluation Matrix},
  author = {{AKI Platform}},
  year = {2026},
  url = {https://aki1k.com/benchmarks/coding}
}