# AKI Capability Intelligence Index™

**Citation Object**: `aki:capabilities:top100` | **Edition**: 2026.4.0 | **Updated**: 2026-10-11

## Top Frontier AI Tools & Autonomous Agents

| Tool / System | Vendor | Model Type | Pricing Model | Est. Cost / 1k Tokens |
| :--- | :--- | :--- | :--- | :---
| **Claude 4 Opus** (v4.1) | Anthropic | Proprietary | subscription | £0.0062 |
| **GPT-6 Omni** (v6.2-prod) | OpenAI | Proprietary | subscription | £0.0052 |
| **Gemini 3 Pro** (v3.1-preview) | Google DeepMind | Proprietary | payg | £0.0026 |
| **Grok 3 Agent** (v3.0-live) | xAI | Proprietary | subscription | £0.0041 |
| **Llama 4 405B Instruct** (v4.0) | Meta Open Source | Open Weights | free | £0.0011 |
| **Claude 3.7 Sonnet Hybrid** (v3.7) | Anthropic | Proprietary | subscription | £0.0048 |
| **GPT-5 Turbo** (v5.1) | OpenAI | Proprietary | payg | £0.0031 |
| **Gemini 2.5 Flash** (v2.5) | Google DeepMind | Proprietary | payg | £0.0007 |
| **DeepSeek V3 MoE** (v3.0) | DeepSeek AI | Open Weights | payg | £0.0005 |
| **Qwen 2.5 Max** (v2.5) | Alibaba Cloud | Proprietary | payg | £0.0018 |
| **Mistral Large 3** (v3.0) | Mistral AI | Proprietary | payg | £0.0035 |
| **Command R+ 2** (v2.0) | Cohere | Proprietary | subscription | £0.0055 |
| **Llama 4 70B** (v4.0) | Meta Open Source | Open Weights | free | £0.0004 |
| **Devin 2.0 SWE Agent** (v2.2) | Cognition AI | Proprietary | subscription | £0.0820 |
| **Cursor Composer v3** (v3.0) | Anysphere | Proprietary | subscription | £0.0125 |
| **Windsurf Cascade Flow** (v2.1) | Codeium | Proprietary | subscription | £0.0098 |
| **Sweep AI Enterprise** (v3.4) | Sweep | Proprietary | subscription | £0.0150 |
| **Aider Architect Pro** (v1.2) | Paul Gauthier | Open Weights | free | £0.0042 |
| **OpenHands Agent v2** (v2.0) | All-Hands AI | Open Weights | free | £0.0045 |
| **Amazon Q Developer Pro** (v2026.3) | AWS | Proprietary | subscription | £0.0078 |
| **Sourcegraph Cody v3** (v3.0) | Sourcegraph | Proprietary | subscription | £0.0080 |
| **Replit Agent Pro** (v2.0) | Replit | Proprietary | subscription | £0.0110 |
| **Cohere Transact** (v1.5) | Cohere | Proprietary | enterprise | £0.0095 |
| **Yi-Lightning Reasoning** (v1.0) | 01.AI | Proprietary | payg | £0.0009 |

## Evaluator Credibility Layer (Meta-Audit)

| Rank | Evaluator Platform | Credibility Score | Transparency | Reproducibility | Independence | Red Flags |
| :--- | :--- | :--- | :--- | :--- | :--- | :---
| #1 | [SWE-bench Verified](https://swebench.com) | **92.6** / 100 | 96% | 94% | 92% | None |
| #2 | [GPQA Diamond](https://arxiv.org/abs/2311.12022) | **88.3** / 100 | 92% | 90% | 94% | None |
| #3 | [Artificial Analysis Index](https://artificialanalysis.ai) | **86.4** / 100 | 88% | 82% | 86% | Ignores human integration effort minutes |
| #4 | [Humanitys Last Exam](https://agi.safe.ai/hle) | **81.8** / 100 | 82% | 78% | 88% | Sample size limited to ~3,000 frontier questions |
| #5 | [LMSYS Chatbot Arena](https://arena.lmsys.org) | **79.5** / 100 | 75% | 68% | 85% | Vulnerable to style/verbosity preference bias over verified task completion |
| #6 | [OSWorld Desktop Bench](https://os-world.github.io) | **74.2** / 100 | 74% | 65% | 82% | High variance across execution order and non-deterministic UI shifts |
| #7 | [MMLU-Pro](https://github.com/TIGER-AI-Lab/MMLU-Pro) | **78.8** / 100 | 86% | 88% | 82% | Near-saturation (top models >90%) limits frontier differentiation |
| #8 | [GAIA Benchmark](https://huggingface.co/spaces/gaia-benchmark/leaderboard) | **80.5** / 100 | 84% | 80% | 85% | Fragile multi-modal file attachments and dead web link dependencies |
| #9 | [SWE-bench Lite](https://swebench.com/lite) | **90.1** / 100 | 94% | 92% | 92% | Reduced 300-issue cohort over-indexes on syntax refactors |
| #10 | [LiveBench AI](https://livebench.ai) | **87.2** / 100 | 89% | 85% | 88% | None |
| #11 | [FrontierMath Epoch](https://epochai.org/frontiermath) | **87.8** / 100 | 90% | 88% | 92% | Extremely low solve rates (<5%) yield wide binomial variance |
| #12 | [ARC Prize AGI](https://arcprize.org) | **89.2** / 100 | 92% | 94% | 90% | None |
| #13 | [Scale AI SEAL Leaderboards](https://scale.com/seal) | **81.6** / 100 | 82% | 74% | 80% | Private test set prevents external third-party reproducibility |
| #14 | [AlpacaEval 2.0](https://github.com/tatsu-lab/alpaca_eval) | **77.8** / 100 | 78% | 84% | 80% | High susceptibility to model length gaming and conversational flattery |
| #15 | [WebArena](https://webarena.dev) | **76.4** / 100 | 76% | 70% | 84% | Complex self-hosted docker harness creates high evaluation friction |
| #16 | [Berkeley Function Calling (BFCL)](https://gorilla.cs.berkeley.edu/leaderboard.html) | **86.5** / 100 | 88% | 86% | 86% | None |

## Empirical Task Vectors & Leading Systems

| Task Vector | Leading System | Top Capability Score | Total Verified Tests |
| :--- | :--- | :--- | :---
| **Autonomous GitHub Issue Resolution** | Devin 2.0 SWE Agent | **95.2** / 100 | 8 tests |
| **OSWorld Computer-Use & Desktop Navigation** | Devin 2.0 SWE Agent | **94.0** / 100 | 8 tests |
| **AIME & GPQA Diamond Multi-Hop Logic** | GPT-6 Omni | **95.0** / 100 | 8 tests |
| **1M+ Token Needles & Contract Extraction** | Gemini 3 Pro | **98.2** / 100 | 8 tests |
| **Multi-Step REST/MCP Tool Chain Execution** | Claude 4 Opus | **96.2** / 100 | 8 tests |
| **Whole-Repository Architectural Migration** | Claude 4 Opus | **95.5** / 100 | 8 tests |
