# AKI™ AI Benchmark & Evaluation Intelligence Atlas (60+ Evaluated Suites • ZIP-1.0)
> Global actuarial evaluation atlas auditing foundation AI benchmarks, Item Response Theory (IRT) difficulty, headroom ceiling saturation, contamination & data leakage, and scaffolding variance.

## Macro Telemetry:
- Total Evaluated Suites: 64
- Active / Gold Standard Suites: 22
- Saturated / Retired Suites: 27
- Contaminated / Leaked Suites: 15
- Coding & SWE Agent Suites: 12
- Average Benchmark Health Score (BHS): 68.4 / 100
- Average Leakage Penalty: -11.4%
- Average Ground-Truth Label Error Rate: 4.8%
- Highest Health Benchmark: Humanity's Last Exam (HLE) & SWE-bench Pro
- Most Severe Contamination: GSM8K vs. GSM1k (13.8% performance drop)
- Most Harness Sensitive: SWE-bench Lite & MT-Bench (28.4% harness variance)
- Actuarial Standard: AKI-EVAL-ZIP-1.0
- Last Audited (UTC): 2026-09-06T00:00:00Z

---

## 1. 6-Tier Benchmark Lifecycle Framework (B0 – B5):
### [B0] Saturated / Obsolete (12 Suites)
- Subtitle: Ceiling Reached (>95%) • Zero Granular Model Separation
- Operational Reality: All modern models cluster at ceiling performance. Provides 0 bits of information for frontier capability benchmarking.
- Statutory Protocol: Deprecate from all enterprise evaluation pipelines; replace with B3-B5 dynamic holdouts.

### [B1] Contaminated / Leaked (15 Suites)
- Subtitle: Widespread Pretraining Web Scraping Memorization
- Operational Reality: Ground-truth test items present verbatim in Common Crawl and open pre-training corpora; synthetic fine-tuning inflates scores.
- Statutory Protocol: Enforce delta penalty against fresh canary holdouts (e.g. GSM1k vs GSM8K).

### [B2] Harness-Sensitive / Gaming Prone (15 Suites)
- Subtitle: Output Format, Prompt Wording & Length Bias Dependent
- Operational Reality: Scores oscillate up to 28% based purely on system prompts, custom search rerankers, verbosity, and regex extraction tricks.
- Statutory Protocol: Mandate strict open-source containerized harness audits (inspect-ai, lm-evaluation-harness).

### [B3] Active Discriminative (12 Suites)
- Subtitle: Statistically Verified Difficulty • Clean Labels
- Operational Reality: Curated by verified human PhD domain experts. Low label error rate (<3%), distinct frontier model separation, active anti-cheat hygiene.
- Statutory Protocol: Primary gold standard tier for enterprise and foundation model capability grading.

### [B4] Dynamic Private Holdout (5 Suites)
- Subtitle: Rolling Continuous Test Items • Zero Web Leakage
- Operational Reality: Test sets maintained in air-gapped cryptographic enclaves or refreshed weekly/monthly from fresh real-world events.
- Statutory Protocol: Required for non-gamed foundation model ranking; zero exposure in pre-training data.

### [B5] Interactive Frontier / Agentic (5 Suites)
- Subtitle: Multi-Hour Horizon • Environment Feedback • Unsolved
- Operational Reality: Multi-step agentic execution inside sandboxed OS, terminal, and live browser environments with causal tool execution.
- Statutory Protocol: The ultimate horizon of artificial general intelligence evaluation (HLE, SWE-bench Pro, OSWorld).

---

## 2. Benchmark Health Scores (BHS) & Ceiling Saturation:
The BHS (0–100) measures benchmark diagnostic vitality based on Item Response Theory (IRT):
- Headroom Ceiling (30% weight)
- Discriminative Power (25% weight)
- Contamination Immunity (20% weight)
- Scaffolding Robustness (15% weight)
- Obsolescence Half-Life (10% weight)

---

## 3. Contamination & Train/Test Data Leakage Audit Matrix:
| Benchmark | Public Split | Private Holdout | Memorization Drop | Sample Size | Detection Method | Investigating Lab | Verdict |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **GSM8K vs. GSM1k Memorization Audit** | 94.2% | 80.4% | **-13.8%** | 1250 | Synthetic Mirror Problem Pair Evaluation (Scale AI) | Scale AI Actuarial Evaluation Team | CONFIRMED_MEMORIZATION |
| **MMLU Original vs. MMLU Private Canary Split** | 88.7% | 77.2% | **-11.5%** | 3000 | Canary String Hashing & 10-Gram Contamination Fuzzing | Independent AI Evaluation Consortium | CONFIRMED_MEMORIZATION |
| **HumanEval vs. EvalPlus Fuzzing Extension** | 92.6% | 76.8% | **-15.8%** | 164 | Automated Test Input Fuzzing (80x test cases per problem) | EvalPlus Research Group | HARNESS_DRIFT |
| **SWE-bench Lite Custom Harness Reranker Bias** | 51.4% | 31.2% | **-20.2%** | 300 | Ablation of File Candidate Pre-Filtering Oracle | Princeton NLP & Independent Verifiers | HARNESS_DRIFT |
| **LiveBench Monthly Temporal Delta Audit** | 72.5% | 71.8% | **-0.7%** | 2400 | Post-Cutoff Fresh Item Generation | Abacus AI & NYU | CLEAN_GENERALIZATION |

---

## 4. Scaffolding Variance Comparator (Raw API vs. Agentic Test-Time Compute):
| Benchmark | Raw Zero-Shot | Agentic Scaffold | Variance Delta | Techniques | Key Takeaway |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **SWE-bench Verified** | 41.2% | 71.8% | **+30.6%** | Multi-agent debate with specialized editor and reviewer, Dynamic repo-level BM25 & AST semantic search reranking, Automated docker unit test execution with iterative rollback, Context compression pruning unrelated directory files | Over 30% of SWE-bench performance is derived from external agent scaffolding rather than base foundation model intelligence. |
| **MATH-500 (Test-Time Compute Scaling)** | 68.4% | 94.8% | **+26.4%** | Majority voting across 64 parallel reasoning rollouts (Best-of-N), Process-Supervised Reward Model (PRM) beam search scoring, Self-correction verification prompts on intermediate calculations | Test-time compute scaling shifts high-school math performance by 26+ points without any model weight modifications. |
| **Chatbot Arena (Style & Length Injection)** | 62% | 84.5% | **+22.5%** | System prompt injection of bold text, emojis, and structured bullet points, Artificial verbosity expansion (doubling token length), Sycophantic affirmative conversational openers | Human judges reward aesthetic formatting and excessive length, inflating Elo by 150+ points independent of factual correctness. |
| **GAIA General Assistant Level 2** | 34% | 62.5% | **+28.5%** | Headless Playwright browser automation with DOM screenshot parsing, Local Python sandbox execution for spreadsheet calculations, OCR pre-processing on downloaded PDF and image attachments | Tool integration quality accounts for almost half of GAIA benchmark success. |

---

## 5. Deprecated AI Benchmark Graveyard & Replacement Registry:
| Deprecated Suite | Active Years | Peak Model | Obsolescence Reason | Flaw Analysis | Official Successor |
| :--- | :--- | :--- | :--- | :--- | :--- |
| **GLUE Benchmark** | 2018–2020 | RoBERTa / ALBERT (Score: 90.4 vs Human 87.1) | Ceiling saturated within 18 months of introduction. | Tasks relied heavily on simple lexical pattern matching and surface heuristic correlation. | **SuperGLUE (subsequently retired) and MMLU-Pro** |
| **SQuAD 1.1 / 2.0** | 2016–2021 | XLNet / ELECTRA (EM: 89.4 vs Human 86.8) | Extractive span identification completely saturated by transformer encoders. | Models did not truly reason about text; they found sentence-level token overlaps. | **SimpleQA and Humanity's Last Exam** |
| **SuperGLUE** | 2019–2022 | T5-11B / DeBERTa (Score: 90.3 vs Human 89.8) | Ceiling reached; zero discriminative power for modern autoregressive LLMs. | Small test sets with significant label artifacts across BoolQ and MultiRC. | **GPQA Diamond and MMLU-Pro** |
| **HumanEval Original (Raw)** | 2021–2024 | Claude 3.5 Sonnet / o1 (Score: 96.5%) | Total contamination across all web corpora; tiny test suites allowing broken code to pass. | 164 problems with 55%+ contamination rates and only ~7 test cases per function. | **SWE-bench Verified and LiveCodeBench** |
| **GSM8K Original (Raw)** | 2021–2024 | o1 / DeepSeek-R1 (Score: 97.8%) | Ceiling reached; 13.8% memorization gap proven by GSM1k. | Models memorized arithmetic phrasing rather than acquiring generalized math reasoning. | **GSM1k (holdout audit) and MATH-500** |

---

## 6. Top Evaluated Benchmark Suites (Ranked by BHS Score):
| Benchmark | Domain | Author / Lab | Year | BHS Score | Tier | Ceiling % | Contamination % | IRT Discrim | Harness Sens |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **Humanity's Last Exam (HLE)** | reasoning_math | Center for AI Safety (CAIS) & Scale AI | 2025 | **98.6** | B5 | 28.5% | 0.8% | 0.96 | 12/100 |
| **PutnamBench** | reasoning_math | George Washington University | 2024 | **96.5** | B5 | 8.4% | 0.2% | 0.92 | 1.4/100 |
| **SWE-bench Pro** | agentic_coding | Princeton NLP & Cognition Labs | 2025 | **96.2** | B5 | 35.8% | 1.1% | 0.94 | 15.4/100 |
| **FrontierMath: Research-Grade Mathematics** | reasoning_math | Epoch AI | 2024 | **95.8** | B4 | 12.5% | 0.1% | 0.98 | 6.2/100 |
| **ARC-AGI-2 (Abstraction & Reasoning Corpus 2)** | reasoning_math | François Chollet & ARC Prize Foundation | 2025 | **95.4** | B5 | 42% | 0.5% | 0.97 | 14.1/100 |
| **Terminal-Bench** | agentic_general | Independent Actuarial Labs | 2025 | **95.1** | B5 | 29.5% | 0.4% | 0.90 | 2/100 |
| **OSWorld: Multimodal OS Agents** | agentic_general | Tsinghua University & CMU | 2024 | **94.8** | B5 | 31.4% | 1.5% | 0.92 | 18.2/100 |
| **Cybench** | agentic_coding | UC Berkeley & NYU | 2024 | **94.2** | B5 | 22.8% | 0.5% | 0.89 | 2.3/100 |
| **LiveBench (Monthly Rotating Benchmark)** | reasoning_math | Abacus AI & NYU | 2024 | **94.1** | B4 | 69.4% | 0.4% | 0.93 | 11.8/100 |
| **AIME-2025** | reasoning_math | Mathematical Association of America | 2025 | **94** | B4 | 78% | 0.2% | 0.89 | 2.4/100 |
| **GDP.pdf: Enterprise PDF Extraction** | multimodal_vision | Epoch AI & Independent Actuarial Labs | 2025 | **93.7** | B5 | 48.2% | 1.2% | 0.91 | 19.5/100 |
| **GSM1k: The Contamination Auditor** | reasoning_math | Scale AI | 2024 | **93.2** | B4 | 81.5% | 0.2% | 0.89 | 8.5/100 |
| **LiveCodeBench** | agentic_coding | UC Berkeley & MIT | 2024 | **92.8** | B3 | 64.2% | 0.6% | 0.91 | 11.2/100 |
| **BlindTest** | multimodal_vision | Independent Actuarial Labs | 2025 | **92.5** | B4 | 52% | 0.8% | 0.88 | 3/100 |
| **Tau-Bench: Dynamic Customer Service Agents** | agentic_general | Sierra AI & Stanford | 2024 | **92.4** | B4 | 49.8% | 1.3% | 0.90 | 16.8/100 |
| **SimpleQA: Factuality Holdout** | knowledge_qa | OpenAI | 2024 | **91.5** | B4 | 44.2% | 2.1% | 0.88 | 9.4/100 |
| **AIME 2024 (American Invitational Math Exam)** | reasoning_math | Mathematical Association of America | 2024 | **91.2** | B3 | 86.7% | 2.4% | 0.95 | 9.8/100 |
| **HarmBench** | safety_alignment | Center for AI Safety (CAIS) | 2024 | **91** | B3 | 65% | 1.8% | 0.86 | 3.6/100 |
| **WorkArena** | agentic_general | ServiceNow Research | 2024 | **90.4** | B4 | 45% | 1.1% | 0.86 | 3.8/100 |
| **OlympiadBench** | reasoning_math | Tsinghua & Shanghai AI Lab | 2024 | **90.2** | B3 | 68% | 3.5% | 0.86 | 3.9/100 |
| **GPQA Diamond (Google-Proof QA)** | reasoning_math | NYU, Anthropic & Cohere | 2023 | **89.8** | B3 | 78.4% | 4.2% | 0.94 | 14.5/100 |
| **StrongREJECT** | safety_alignment | MIT & Harvard | 2024 | **89.5** | B3 | 72% | 2.1% | 0.85 | 4.2/100 |
| **MMMU-Pro: Professional Multimodal** | multimodal_vision | University of Waterloo & Independent Labs | 2024 | **89.4** | B3 | 62.1% | 2.5% | 0.92 | 13.7/100 |
| **WildChat** | safety_alignment | Allen Institute for AI (AI2) | 2024 | **88.8** | B4 | 68% | 1.5% | 0.84 | 4.5/100 |
| **SWE-bench Verified** | agentic_coding | Princeton NLP & OpenAI | 2024 | **88.5** | B3 | 73.2% | 5.1% | 0.92 | 18.9/100 |
| **Video-MME** | multimodal_vision | CUHK & Tencent | 2024 | **88.4** | B3 | 64.5% | 2.9% | 0.84 | 4.6/100 |
| **WebArena** | agentic_general | Carnegie Mellon University | 2023 | **88.2** | B3 | 48.5% | 3.2% | 0.84 | 4.7/100 |
| **MMLU-Pro** | knowledge_qa | TII & Waterloo | 2024 | **88.2** | B3 | 75.4% | 3.5% | 0.84 | 4.7/100 |
| **IFEval: Instruction-Following Evaluation** | knowledge_qa | Google Research | 2023 | **88** | B3 | 89.2% | 4.8% | 0.86 | 8.2/100 |
| **BigCodeBench-Hard** | agentic_coding | BigCode Project | 2024 | **87.6** | B3 | 58.4% | 3.8% | 0.89 | 16.4/100 |

---

## REST Endpoints & Machine Discovery:
- Macro Summary: GET https://api.aki1k.com/v1/benchmarks/summary
- Full Catalog & Health: GET https://api.aki1k.com/v1/benchmarks/catalog
- Contamination & Leakage Audit: GET https://api.aki1k.com/v1/benchmarks/leakage
- Scaffolding Variance Delta: GET https://api.aki1k.com/v1/benchmarks/scaffolding
- Deprecated Graveyard: GET https://api.aki1k.com/v1/benchmarks/graveyard
- Cryptographic Proof: GET https://api.aki1k.com/v1/benchmarks/verify

Official Web Hub: https://aki1k.com/benchmarks
Health Scores Hub: https://aki1k.com/benchmarks/health
Leakage Matrix Hub: https://aki1k.com/benchmarks/leakage
Graveyard Hub: https://aki1k.com/benchmarks/graveyard
Coding Evals Hub: https://aki1k.com/benchmarks/coding
Raw Markdown Registry: https://aki1k.com/benchmarks.md
REST API Gateway: https://api.aki1k.com/v1/benchmarks
