# AKI Grounded Intelligence Engine™ v2.2 Top 50 Forward-Looking Anti-Softness Release

> **Standard:** Zero-Incentive-Protocol (ZIP-1.0) • Independent Empirical Consensus
> **Snapshot Date:** 2026-10-03 • **Edition:** October 2026 Production Release
> **Leader:** Gemini 4 Argon High (93.2/100, Conf 88%) • **Top Open:** Llama Spark 1.3 Max (88.4/100)

## Macro Telemetry Summary

- **Total Models Audited:** 50 Frontier & Mid-Tier Cognitive Systems
- **Mean AKI Score:** 75/100
- **Mean Evidence Confidence:** 86%
- **Open Weights Parity:** 22 Open (44%) vs 28 Proprietary
- **Fastest Trajectory (90-Day Improvement):** Grok 4.7 xhigh (+3.00 velocity bonus)
- **Efficiency Champions:** Qwen 3.8 Max ($0.40 / $0.80) & DeepSeek V4.1 Flash Max ($0.20 / $0.80)

## v2.2 Anti-Softness Weighting

| Dimension | Weight | Primary Benchmarks & Role |
| :--- | :---: | :--- |
| Hard Capability | 29% | AA Private Set, HLE, CritPt, FrontierMath, SWE-Pro |
| Coding & SWE | 18% | SWE-bench Verified/Pro, LiveCodeBench, SciCode |
| Reasoning & Math | 14% | GPQA Diamond, HLE, Formal Theorem Proving |
| Agentic / Computer Use | 13% | OSWorld 2.0, Multi-turn Tool Chaining (Up from 11%) |
| Science & Knowledge | 10% | Scientific reasoning, non-hallucination verification |
| Efficiency & Value | 7% | Cost-per-task, tokens/sec context-window economics |
| Openness / Reproducibility | 5% | Open weights (+5), open API (+2) (Up from 3%) |
| Freshness & Stability | 3% | Evidence recency, low day-to-day volatility |
| Preference Diversity | 1% | Blind human Elo (heavily discounted, vote-skew penalized) |

## Anti-Hype Adjustments Applied After Base Score

- **Velocity Bonus:** `min(3.0, 90d_improvement * 0.6)` (Rewards rising architectures permanently)
- **Social Velocity Penalty:** `max(0, (Weighted_Social_Share - 20%) * 0.5)` (Deducts points if social share > 20%)
- **Media Hype Penalty:** Gaps between vendor announcements and independent replications penalized up to -5 pts
- **Single Source Cap:** No evaluator family may exceed 22% of final score
- **Lab Concentration Cap:** No single lab may hold >3 of Top 5 positions

## Top 50 Intelligence Master Ledger

| Rank | Model Name | Provider | License | AKI Score | Conf % | Spread | Pricing (1M in/out) | Change | Why Ranked Here (Deterministic Synthesis) |
| :---: | :--- | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :--- |
| #1 | **Gemini 4 Argon High** | Google | Proprietary | **93.2** | 88% | ±1.8 | $2.00 / $10.00 | +2.1 | Why #1: Strong early independent signals + Arena + high velocity — Top drivers: Hard Capability 93rd percentile + Agentic 94th percentile — Rate of improvement +4.8 last 90 days → velocity bonus +2.88 — Social velocity penalty 0 — Pricing $2/$10 excellent value — Forward-looking #1. |
| #2 | **GPT-6 Astra Max** | OpenAI | Proprietary | **92** | 93% | ±1.1 | $10.00 / $50.00 | -0.4 | Why #2: Extremely broad independent evidence — Top drivers: Hard Capability 96th percentile + Reasoning 95th percentile — Social velocity penalty -5 applied — Pricing $10/$50 — Highest multi-source consensus — Slightly reduced historical advantage Hard 32%→29%. |
| #3 | **Claude Opus 5.5 Max/Thinking** | Anthropic | Proprietary | **90.1** | 93% | ±1.1 | $4.00 / $20.00 | -0.5 | Why #3: Deep independent evidence base — Top drivers: Coding & SWE 94th percentile + Hard 92nd percentile — Social velocity penalty -7.5 applied — Still elite hype filtered — Pricing $4/$20 — Confidence 95%→93% reduced historical advantage. |
| #4 | **Llama Spark 1.3 Max** | Meta | Open | **88.4** | 90% | ±1.5 | $1.25 / $4.25 | +2.1 | Why #4: Solid agentic + multimodal + velocity — Top drivers: Agentic 91st percentile + Openness 95th percentile open-weight — Velocity +1.86 — Agentic 11%→13% +1.0 + openness 3%→5% +1.0 — Pricing ~$1.25/$4.25 good price — Rising. |
| #5 | **Claude Sonnet 5.5 Max** | Anthropic | Proprietary | **88.2** | 91% | ±1.4 | $2.00 / $10.00 | -1.5 | Why #5: Excellent breadth + value — Top drivers: Coding 90th percentile + Efficiency 89th percentile — Velocity +0.4 — Best pure value top tier — Outstanding coding/agent + value less hype than Opus/Fable. |
| #6 | **Grok 4.7 xhigh** | xAI | Proprietary | **88** | 88% | ±1.8 | $2.00 / $6.00 | +2.9 | Why #6: Clear agentic & coding progress + velocity — Top drivers: Agentic 93rd percentile + Coding 87th percentile — Velocity +3 cap fastest improvement +5.2 last 90 days — Pricing $2/$6 — X-native tools. |
| #7 | **Qwen 3.8 Max** | Alibaba | Open | **86.1** | 90% | ±1.5 | $0.40 / $0.80 | +2.5 | Why #7: Strong coding + multilingual + open — Top drivers: Coding 89th percentile + Openness 94th percentile open-weight — Velocity +1.68 — Openness +1.0 — Leading open-weight. |
| #8 | **DeepSeek V4 Pro Max** | DeepSeek | Open | **85.2** | 90% | ±1.5 | $0.66 / $1.98 | +3.1 | Why #8: Strong open-weight reasoning + value + velocity — Top drivers: Efficiency 92nd percentile best pure value + Openness 95th percentile — Velocity +1.44 — Extreme cost efficiency + open. |
| #9 | **GLM-5.3 Max** | Z.AI | Open | **85** | 89% | ±1.6 | $1.40 / $4.40 | +2.1 | Why #9: Competitive reasoning/coding + open — Top drivers: Reasoning 87th percentile + Openness 93rd percentile — Openness +1.0 — Good value. |
| #10 | **Kimi K3 Max** | Moonshot | Proprietary | **84.8** | 90% | ±1.5 | $2.70 / $13.50 | +0.6 | Why #10: Exceptional long-context + velocity — Top drivers: Reasoning 88th percentile + Agentic 86th percentile long-context — Velocity +0.6 — Solid strong reasoning & long context. |
| #11 | **GPT-6.1 Sol Max** | OpenAI | Proprietary | **84.5** | 87% | ±1.9 | $2.00 / $10.00 | -4.4 | Why #11: Near-Astra intelligence at low cost — Top drivers: Efficiency 91st percentile + Hard 85th percentile — Near-Astra at ~1/5 cost — Efficiency king — Was 88.9 #5 now #11 because hard reduction hits efficiency-focused model. |
| #12 | **Claude Fable 5.1 Max** | Anthropic | Proprietary | **84.3** | 91% | ±1.4 | $10.00 / $50.00 | -3.5 | Why #12: Strong coding & long-horizon — Top drivers: Coding 88th percentile + Agentic 87th percentile long-horizon — Strong coding & long-horizon — Still top-tier after penalty expensive + high social noise. |
| #13 | **MiMo-V2.6-Pro** | Xiaomi | Open | **83.2** | 87% | ±1.9 | $0.40 / $0.80 | +3.3 | Why #13: Leading open-weight contender + velocity — Top drivers: Openness 96th percentile open-weight + Efficiency 90th percentile $0.40/$0.80 best value — Velocity +1.2 — Best open-weight overall. |
| #14 | **Claude Opus 5** | Anthropic | Proprietary | **80.9** | 89% | ±1.6 | $5.00 / $25.00 | -0.5 | Why #14: Deep historical evidence — Top drivers: Hard 84th percentile + Coding 83rd percentile — Previous-gen but still competitive — Deep historical evidence. |
| #15 | **GPT-6 Sol Max** | OpenAI | Proprietary | **80.2** | 85% | ±2.2 | $2.00 / $10.00 | -0.5 | Why #15: Balanced everyday frontier — Top drivers: Efficiency 86th percentile + Reasoning 82nd percentile — Reliable — Balanced frontier everyday model. |
| #16 | **DeepSeek V4.1 Flash Max** | DeepSeek | Open | **80.5** | 89% | ±1.6 | $0.20 / $0.80 | +1.3 | Why #16: Extreme efficiency + open — Top drivers: Efficiency 94th percentile budget king + Openness 94th percentile — Openness + velocity — Budget king cheapest high-performer. |
| #17 | **Gemini 3.8 Flash** | Google | Proprietary | **78.9** | 88% | ±1.8 | $0.75 / $3.75 | +0.5 | Why #17: Speed + price — Top drivers: Efficiency 88th percentile + Agentic 82nd percentile — High-volume workhorse. |
| #18 | **Step 5 Preview** | StepFun | Proprietary | **78.8** | 84% | ±2.4 | $0.50 / $2.00 | +1.2 | Why #18: Emerging frontier — Top drivers: Reasoning 82nd percentile + Hard 80th percentile — Velocity +1.2 — Emerging frontier. |
| #19 | **Grok 4.6** | xAI | Proprietary | **77.3** | 85% | ±2.2 | $2.00 / $6.00 | +0.5 | Why #19: Previous generation — Top drivers: Coding 81st percentile + Agentic 80th percentile — Previous Grok generation. |
| #20 | **Claude Fable 5** | Anthropic | Proprietary | **75.8** | 89% | ±1.6 | $10.00 / $50.00 | -0.3 | Why #20: Deep evidence base — Top drivers: Coding 82nd percentile + Reasoning 81st percentile — Deep evidence base. |
| #21 | **Qwen 3.8 27B** | Alibaba | Open | **77** | 87% | ±1.9 | $0.20 / $0.60 | +1.7 | Why #21: Strong open model + open — Top drivers: Openness 93rd percentile + Coding 84th percentile — Strong open model. |
| #22 | **GLM-5.3 Flash** | Z.AI | Open | **76.2** | 86% | ±2.1 | $0.15 / $0.50 | +1.5 | Why #22: High efficiency open + open — Top drivers: Efficiency 92nd percentile extreme value + Openness 92nd percentile — Extreme value. |
| #23 | **DeepSeek V4 Flash** | DeepSeek | Open | **75.8** | 85% | ±2.2 | $0.14 / $0.28 | +1.7 | Why #23: Budget high-performer + open — Top drivers: Efficiency 93rd percentile + Openness 93rd percentile — Budget high-performer. |
| #24 | **GPT-6 Luna** | OpenAI | Proprietary | **72.8** | 80% | ±3 | $0.10 / $0.50 | -0.4 | Why #24: Ultra-cheap tier — Top drivers: Efficiency 90th percentile ultra-cheap + Freshness 80th percentile — High-volume only not intelligence-focused. |
| #25 | **MiniMax M3** | MiniMax | Open | **74** | 81% | ±2.8 | $0.30 / $1.20 | +1.5 | Why #25: Growing evidence + open — Top drivers: Openness 88th percentile + Reasoning 78th percentile — Growing evidence — Emerging. |
| #26 | **GPT-6 Mini** | OpenAI | Proprietary | **73.5** | 82% | ±2.7 | $0.20 / $1.00 | NEW | Why #26: New entry #26 — GPT-6 Mini — AKI 73.5 · Conf 82% — Near-GPT-6 Sol intelligence at mini cost — Top drivers: Efficiency 91st percentile + Reasoning 79th percentile — Freshness boost +8% new high-quality independent evidence <14 days. |
| #27 | **Claude Haiku 5.5** | Anthropic | Proprietary | **73** | 84% | ±2.4 | $0.50 / $2.50 | NEW | Why #27: New entry #27 — Claude Haiku 5.5 — AKI 73.0 · Conf 84% — Fast + cheap Claude — Top drivers: Efficiency 89th percentile + Coding 80th percentile — Social velocity penalty -3 — Pricing $0.50/$2.50 — Emerging. |
| #28 | **Gemini 4 Flash High** | Google | Proprietary | **72.8** | 83% | ±2.5 | $0.50 / $2.00 | NEW | Why #28: New entry #28 — Gemini 4 Flash High — AKI 72.8 · Conf 83% — Speed + Argon lineage — Top drivers: Efficiency 90th percentile + Agentic 83rd percentile — Freshness boost +10% — Pricing $0.50/$2.00 — Emerging. |
| #29 | **Kimi K4 Max** | Moonshot | Proprietary | **71.8** | 82% | ±2.7 | $2.50 / $12.00 | NEW | Why #29: New entry #29 — Kimi K4 Max — AKI 71.8 · Conf 82% — Exceptional long-context 1M+ — Top drivers: Reasoning 84th percentile + Openness 90th percentile — Freshness boost +8% new high-quality independent evidence <14 days. |
| #30 | **DeepSeek V5 Pro** | DeepSeek | Open | **71.5** | 83% | ±2.5 | $0.70 / $2.10 | NEW | Why #30: New entry #30 — DeepSeek V5 Pro — AKI 71.5 · Conf 83% — Next-gen open-weight — Top drivers: Openness 95th percentile + Coding 82nd percentile — Velocity +1.5 — Pricing ~$0.70/$2.10 — Emerging. |
| #31 | **Qwen 3.9 Max** | Alibaba | Open | **70.8** | 84% | ±2.4 | $2.20 / $6.50 | NEW | Why #31: New entry #31 — Qwen 3.9 Max — AKI 70.8 · Conf 84% — Coding + multilingual + open — Top drivers: Coding 85th percentile + Openness 94th percentile — Velocity +1.2 — Pricing $2.20/$6.50 — Emerging. |
| #32 | **GLM-6 Preview** | Z.AI | Open | **70.2** | 81% | ±2.8 | $1.50 / $4.80 | NEW | Why #32: New entry #32 — GLM-6 Preview — AKI 70.2 · Conf 81% — Next-gen reasoning + open — Top drivers: Reasoning 83rd percentile + Openness 92nd percentile — Freshness boost +12% — Pricing ~$1.50/$4.80 — Emerging. |
| #33 | **MiMo-V3 Preview** | Xiaomi | Open | **69.8** | 80% | ±3 | $0.50 / $1.00 | NEW | Why #33: New entry #33 — MiMo-V3 Preview — AKI 69.8 · Conf 80% — Next open-weight contender — Top drivers: Openness 95th percentile + Efficiency 88th percentile — Velocity +1.0 — Pricing ~$0.50/$1.00 — Emerging. |
| #34 | **Llama Spark 1.2** | Meta | Open | **69.5** | 86% | ±2.1 | $1.00 / $3.50 | — | Why #34: Previous Llama — Top drivers: Agentic 82nd percentile + Openness 90th percentile — Previous-gen but still solid agentic + multimodal — Deep historical evidence. |
| #35 | **Grok 4.5** | xAI | Proprietary | **69** | 83% | ±2.5 | $2.00 / $6.00 | — | Why #35: Previous Grok — Top drivers: Coding 79th percentile + Agentic 78th percentile — Previous generation. |
| #36 | **Claude Opus 4.8** | Anthropic | Proprietary | **68.5** | 88% | ±1.8 | $5.00 / $25.00 | — | Why #36: Legacy but still used — Top drivers: Hard 78th percentile + Coding 77th percentile — Deep historical evidence — Legacy. |
| #37 | **GPT-4o / GPT-5 Legacy** | OpenAI | Proprietary | **68** | 90% | ±1.5 | $2.50 / $10.00 | — | Why #37: Legacy frontier — Top drivers: Hard 77th percentile + Reasoning 76th percentile — Deep historical evidence — Legacy but still used. |
| #38 | **Gemini 2.5 Pro** | Google | Proprietary | **67.5** | 89% | ±1.6 | $1.25 / $5.00 | — | Why #38: Legacy Gemini — Top drivers: Reasoning 77th percentile + Efficiency 76th percentile — Deep historical evidence. |
| #39 | **Kimi K2.x** | Moonshot | Proprietary | **67** | 82% | ±2.7 | $1.20 / $4.80 | — | Why #39: Previous Kimi — Top drivers: Reasoning 76th percentile + Agentic 75th percentile — Previous generation. |
| #40 | **Qwen 3.5 Max** | Alibaba | Open | **66.5** | 85% | ±2.2 | $0.80 / $2.40 | — | Why #40: Previous Qwen — Top drivers: Coding 77th percentile + Openness 90th percentile — Previous-gen. |
| #41 | **DeepSeek V3** | DeepSeek | Open | **66** | 87% | ±1.9 | $0.14 / $0.28 | — | Why #41: Previous DeepSeek — Top drivers: Efficiency 87th percentile + Openness 92nd percentile — Previous-gen — Extreme cost efficiency. |
| #42 | **Mistral Large 3** | Mistral | Open | **65.5** | 83% | ±2.5 | $2.00 / $6.00 | — | Why #42: European open-weight — Top drivers: Openness 91st percentile + Reasoning 74th percentile — Competitive open-weight. |
| #43 | **Command R+ / Cohere** | Cohere | Proprietary | **65** | 82% | ±2.7 | $3.00 / $15.00 | — | Why #43: RAG + enterprise — Top drivers: Science & Knowledge 76th percentile + Agentic 74th percentile — Enterprise-focused. |
| #44 | **Claude Sonnet 4.5** | Anthropic | Proprietary | **64.5** | 87% | ±1.9 | $3.00 / $15.00 | — | Why #44: Previous Sonnet — Top drivers: Coding 75th percentile + Efficiency 73rd percentile — Previous-gen. |
| #45 | **GPT-6 Nano** | OpenAI | Proprietary | **64** | 80% | ±3 | $0.05 / $0.20 | NEW | Why #45: New entry #45 — GPT-6 Nano — AKI 64.0 · Conf 80% — Ultra-cheap nano — Top drivers: Efficiency 92nd percentile $0.05/$0.20 — Freshness boost +10% — High-volume only not intelligence-focused. |
| #46 | **Gemini 3.5 Flash** | Google | Proprietary | **63.5** | 86% | ±2.1 | $0.50 / $2.00 | — | Why #46: Previous Flash — Top drivers: Efficiency 85th percentile + Speed — High-volume workhorse — Previous-gen. |
| #47 | **Phi-4 / Microsoft** | Microsoft | Open | **63** | 84% | ±2.4 | $0.10 / $0.30 | — | Why #47: Small open model — Top drivers: Openness 90th percentile + Efficiency 84th percentile — Small but efficient — Open. |
| #48 | **Llama 4 Scout** | Meta | Open | **62.5** | 85% | ±2.2 | $0.80 / $3.00 | — | Why #48: Previous Llama scout — Top drivers: Agentic 76th percentile + Openness 89th percentile — Previous-gen. |
| #49 | **Gemma 3 27B** | Google | Open | **62** | 83% | ±2.5 | $0.15 / $0.40 | — | Why #49: Small open Google — Top drivers: Openness 91st percentile + Efficiency 82nd percentile — Small open-weight — Good value. |
| #50 | **MiniMax M2** | MiniMax | Open | **61.5** | 79% | ±3.1 | $0.20 / $0.80 | — | Why #50: Previous MiniMax — Top drivers: Openness 87th percentile + Reasoning 71st percentile — Growing evidence — Previous-gen — 15 slots reserved for next-gen entries. |

---
*Cryptographically attested by AKI Platform under Zero-Incentive-Protocol (ZIP-1.0). Live API: https://api.aki1k.com/v1/models/top50*
