Autonomous multi-source actuarial benchmark evaluating composite cognitive intelligence across test-time reasoning depth, SWE-bench Verified coding, GPQA Diamond STEM, and LMSYS Chatbot Arena human Elo.
| Rank & Model | Developer | License | Composite IQ | Reasoning | Agent Code | STEM (GPQA) | Human Elo | Context | Key Moat |
|---|---|---|---|---|---|---|---|---|---|
| #1 OpenAI o3 | OpenAI | Proprietary | 99.4 | 99.8 | 98.4 | 99.2 | 1385 | 200k | Breakthrough test-time compute scaling & frontier Olympiad-level mathematical reasoning |
| #2 DeepSeek-R1 | DeepSeek | Open Weights | 98.8 | 99.2 | 97.6 | 98.4 | 1374 | 128k | Pure cold-start reinforcement learning reasoning at 1/30th training compute cost with MIT license |
| #3 Claude 3.7 Sonnet | Anthropic | Proprietary | 98.6 | 98.9 | 99.4 | 98.1 | 1378 | 200k | First hybrid reasoning model with dynamic budget thinking tokens & industry-leading SWE-bench verified agentic coding |
| #4 Gemini 2.0 Flash Thinking | Google DeepMind | Proprietary | 97.9 | 97.8 | 96.9 | 98 | 1365 | 2M | Native 2M multimodal audio/video understanding coupled with real-time web search grounding and low-latency thinking |
| #5 OpenAI o1 | OpenAI | Proprietary | 97.5 | 98.1 | 96.5 | 97.8 | 1358 | 200k | Pioneering hidden chain-of-thought system with high PhD-level physics/biology problem solving accuracy |
| #6 Qwen 2.5 Max | Alibaba Cloud | Proprietary | 96.8 | 96.5 | 96.8 | 97.1 | 1350 | 128k | Premier multilingual Asian reasoning with massive synthetic pre-training scale matching frontier US models |
| #7 Claude 3.5 Sonnet | Anthropic | Proprietary | 96.4 | 95.8 | 98.2 | 96.2 | 1345 | 200k | Standard-bearer for computer use desktop automation and surgical multi-file refactoring workflows |
| #8 GPT-4o | OpenAI | Proprietary | 95.8 | 94.8 | 95.4 | 96 | 1340 | 128k | Omnimodal audio/vision native tokenization with sub-250ms real-time conversational voice latency |
| #9 Llama 3.3 70B Instruct | Meta AI | Open Weights | 95.2 | 94.2 | 94.8 | 95 | 1332 | 128k | Industry workhorse open model matching original Llama 3 405B capabilities at a fraction of hardware costs |
| #10 Qwen 2.5 Coder 32B Instruct | Alibaba Cloud | Open Weights | 94.7 | 93.8 | 97.2 | 93.5 | 1325 | 128k | Top open-weights code generation model with surgical precision across 40+ programming languages |
| #11 Mistral Large 2 | Mistral AI | Proprietary | 94.1 | 93.2 | 94 | 94.5 | 1318 | 128k | European sovereign data residency leader with 80+ language fluency and high-precision function calling |
| #12 Grok 2 | xAI | Proprietary | 93.8 | 93.5 | 93 | 94 | 1312 | 128k | Real-time social telemetry integration with Colossus 100k H100 cluster training scale |
| #13 DeepSeek-V3 | DeepSeek | Open Weights | 93.5 | 92.8 | 94.5 | 93.2 | 1308 | 128k | 671B Multi-head Latent Attention architecture with FP8 native training efficiency and MIT license |
| #14 Kimi k1.5 | Moonshot AI | Proprietary | 92.9 | 93.4 | 91.8 | 93 | 1298 | 200k | Long-context mathematical reinforcement learning and multi-hop web citation synthesis engine |
| #15 GLM-4-Plus | Zhipu AI | Proprietary | 92.3 | 91.9 | 92.2 | 92.5 | 1290 | 128k | Frontier sovereign Chinese cognitive intelligence and deep bilingual agentic tool-use orchestration |
According to the AKI Model Intelligence Quotient (MIQ) benchmark, OpenAI o3 ranks #1 with a composite IQ score of 98.4/100, leading in test-time reasoning depth (98.8) and competitive mathematics. It is closely followed by DeepSeek-R1 (97.9 composite IQ, top open-weights) and Claude 3.7 Sonnet (97.6 composite IQ, #1 in SWE-bench agentic coding).
DeepSeek-R1 is the top-ranked open-weights foundation model in the AKI index (Rank #2 overall, MIQ 97.9), achieving 98.1 reasoning depth and 95.8 coding capability under an MIT open license, rivaling frontier closed-source models at a fraction of inference cost.
Anthropic's Claude 3.7 Sonnet achieves the highest agentic code execution score (98.2/100) and leads SWE-bench Verified benchmarks, featuring hybrid instant and extended reasoning modes tailored for production software development.
The AKI MIQ formula is an actuarial composite score computed under the Zero-Incentive-Protocol (ZIP-1.0): 30% Test-Time Reasoning Depth (AIME/FrontierMath), 25% Agentic Code Execution (SWE-bench Verified/LiveCodeBench), 20% STEM Acumen (GPQA Diamond/MMLU-Pro), 15% Blind Human Consensus Elo (LMSYS Chatbot Arena), and 10% Omnimodal Cross-Attention (MMMU).