Full audit of SWE-bench Verified, SWE-bench Pro, SWE-bench Multimodal, RepoQA, and LiveCodeBench measuring real developer agent autonomy and test pass rates.
Evaluated across Item Response Theory (IRT) Headroom, SOTA Ceilings, Contamination Risk, and Scaffolding Delta.
| Benchmark Suite | Domain | BHS Score | Tier | Ceiling % | Contamination % | IRT Discrim | Harness Sens |
|---|---|---|---|---|---|---|---|
| SWE-bench Pro Princeton NLP & Cognition Labs • 2025 | agentic_coding | 96.2 | B5 | 35.8% | 1.1% | 0.94 | 15.4/100 |
| SWE-bench Verified Princeton NLP & OpenAI • 2024 | agentic_coding | 88.5 | B3 | 73.2% | 5.1% | 0.92 | 18.9/100 |
| LiveCodeBench UC Berkeley & MIT • 2024 | agentic_coding | 92.8 | B3 | 64.2% | 0.6% | 0.91 | 11.2/100 |
| BigCodeBench-Hard BigCode Project • 2024 | agentic_coding | 87.6 | B3 | 58.4% | 3.8% | 0.89 | 16.4/100 |
| SWE-bench Lite Princeton NLP • 2024 | agentic_coding | 62.4 | B2 | 54% | 12.8% | 0.76 | 28.4/100 |
| EvalPlus (HumanEval+) EvalPlus Team • 2023 | agentic_coding | 65.2 | B2 | 92.4% | 16.5% | 0.75 | 14.8/100 |
| HumanEval (Original 164 Problems) OpenAI (Chen et al.) • 2021 | agentic_coding | 24.5 | B0 | 96.5% | 55% | 0.28 | 14.2/100 |
| MBPP (Mostly Basic Python Problems) Google Research (Austin et al.) • 2021 | agentic_coding | 28.6 | B0 | 95.2% | 48.5% | 0.32 | 11/100 |
| Cybench UC Berkeley & NYU • 2024 | agentic_coding | 94.2 | B5 | 22.8% | 0.5% | 0.89 | 2.3/100 |
| InterCode Princeton NLP • 2023 | agentic_coding | 84.5 | B3 | 72% | 4.5% | 0.80 | 6.2/100 |
| RepoBench HKUST • 2023 | agentic_coding | 74 | B2 | 84.5% | 9.8% | 0.70 | 10.4/100 |
| DeepBench Baidu Research • 2017 | agentic_coding | 11 | B0 | 99.5% | 80% | 0.10 | 35.6/100 |
Licensed under CC BY 4.0. Permitted for academic, enterprise, and search engine citation.