Daily mathematical benchmarking of top 100 artificial intelligence models across 26 independent evaluation feeds with single-source caps and 48h freeze windows.
GPT-6 Astra Max by OpenAI leads the AKI Superintelligence Index with an actuarial score of 92.4 (95% CI ±0.3). It commands the frontier across FrontierMath, HLE Diamond (60.6%), and ExploitBench with an unprecedented 4096K context window.
Scores are bound by ZIP-1.0 rules: 12% single-source caps on intelligence benchmarks, 50% max score on un-replicated vendor claims, and automated penalties (-0.3 to -1.6 pts) for social-hype divergence.
MiMo-V2.6-Pro (AKI 86.5, Value Index 98.6) at $0.60/$2.40 per M tokens and DeepSeek V4.1 Flash (AKI 86.2, Value Index 91.2) offer the highest autonomy-to-cost ratios in enterprise production.
Actuarial ranking audited across 26 independent benchmarks. Full 100-model ledger available via API JSON.
| # | Model | Provider | AKI Score | AA Intel | 95% CI | Price In/Out | Openness | Key Strength | Value Idx | Context | Δ prev | Tier |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| #1 |
2026-09-18
|
92.4 | 52.7 | ±0.3 | $10.00 / $50.00 | Proprietary | 42.1 | 4096K | +0.4 | T1 | ||
| #2 |
2026-08-12
|
91.7 | 52.6 | ±0.4 | $3.50 / $12.00 | Proprietary | 78.4 | 2000K | +0.3 | T1 | ||
| #3 |
2026-07-29
|
90.2 | 57.6 | ±0.5 | $4.00 / $20.00 | Proprietary | 38.2 | 1000K | -0.5 | T1 | ||
| #4 |
2026-06-15
|
89.9 | 43.6 | ±0.6 | $2.20 / $8.80 | Proprietary | 71.3 | 1000K | +0.2 | T1 | ||
| #5 |
2026-08-01
|
89.1 | 53.4 | ±0.5 | $1.80 / $7.20 | Open Weights | 82.7 | 1000K | +0.8 | T1 | ||
| #6 |
2026-08-25
|
88.9 | 51.8 | ±0.4 | $2.00 / $10.00 | Proprietary | 75.6 | 1100K | +0.6 | T1 | ||
| #7 |
2026-07-29
|
88.8 | 56.0 | ±0.4 | $2.00 / $10.00 | Proprietary | 68.9 | 1000K | -0.3 | T1 | ||
| #8 |
2026-07-10
|
88.5 | 50.2 | ±0.5 | $0.43 / $0.87 | Open Weights | 87.2 | 1000K | +0.4 | T1 | ||
| #9 |
2026-08-30
|
88.1 | 48.1 | ±0.7 | $0.92 / $3.50 | Open Weights | 90.1 | 262K | +1.5 | T1 | ||
| #10 |
2026-06-20
|
87.8 | 53.4 | ±0.6 | $10.00 / $50.00 | Proprietary | 62.4 | 1000K | -0.1 | T1 | ||
| #11 |
2026-07-15
|
87.4 | 44.8 | ±0.5 | $0.90 / $2.86 | Open Weights | 84.3 | 1000K | +0.5 | T2 | ||
| #12 |
2026-08-05
|
86.5 | 46.3 | ±0.6 | $0.60 / $2.40 | Open Weights | 98.6 | 131K | +0.7 | T2 | ||
| #13 |
2026-07-18
|
86.2 | 39.5 | ±0.4 | $0.09 / $0.18 | Open Weights | 91.2 | 1000K | +1.2 | T2 | ||
| #14 |
2026-06-25
|
85.5 | 40.9 | ±0.3 | $0.75 / $3.75 | Proprietary | 88.9 | 1048K | +1.0 | T2 | ||
| #15 |
2026-08-10
|
85.1 | 46.4 | ±0.7 | $2.00 / $6.00 | Proprietary | 54.2 | 2000K | +0.1 | T2 |
As of 2026, GPT-6 Astra Max (OpenAI) leads the AKI Superintelligence Index with an overall capability rating of 92.4/100, commanding the frontier across FrontierMath, HLE Diamond (60.6%), and ExploitBench with an unprecedented 4096K context window. Gemini 4 Argon High (91.7) and Claude Opus 5.5 Max (90.2) follow closely in Tier 1: Frontier ASI/AGI.
Claude Opus 5.5 Max (Anthropic) is the empirical leader for autonomous coding loops and long agentic workflows, scoring 90.2 AKI with an Artificial Analysis Intelligence rating of 57.6 and SWE-bench Pro V2 accuracy of 99.4%. It excels in extended agentic coherence and zero-guidance repo engineering.
MiMo-V2.6-Pro (Xiaomi) achieves the highest Value Index on the entire ledger at 98.6 (86.5 AKI at $0.60/$2.40 per million tokens), followed by DeepSeek V4.1 Flash at 91.2 Value Index (86.2 AKI at $0.09/$0.18 per million tokens), which commands global developer volume across 33.6T tokens.
Under strict ZIP-1.0 mathematical governance and capability progression, legacy 2024 and 2025 frontier models have been relegated to Tier 4 (Cost Optimized) and Tier 5 (Historical Baselines). Claude 3.7 Sonnet is positioned at #57 (60.0 AKI), OpenAI o3 at #58 (59.6 AKI), and GPT-4o at #64 (57.2 AKI), serving as permanent empirical baselines.
The index enforces strict ZIP-1.0 mathematical governance: no single benchmark may contribute more than 12% of the Intelligence dimension; vendor self-reported figures are capped at 50% max score unless replicated by at least two independent evaluators; new models enter a mandatory 48-hour freeze window; and hype divergence penalties (-0.3 to -1.6 points) are deducted when social media attention deviates significantly from measured capability.