Reasoning leaderboard
Best LLMs for reasoning
Reasoning-heavy tasks — graduate science QA, abstract puzzles, novel logic — separate frontier models from the pack. This ranking weights GPQA Diamond (30%), ARC-AGI-2 (25%), LiveBench (20%), MMLU-Pro (15%) and AA Index (10%).
Category leader
Claude 4.5 Opus.
Strong coding + long-horizon reasoning.
+1.1 vs #2
Ranking
N=48| # | Model | Composite | GPQA | ARC-AGI | LiveBench | MMLU-Pro | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | Claude 4.5 OpusAnthropic | 84.8 | 81.9 | — | — | 90.6 | — |
| 02 | GPT-5.6 TerraOpenAI | 83.7 | 80.2 | — | — | 90.8 | — |
| 03 | Grok 4xAI | 78.5 | 74.1 | — | — | 87.4 | — |
| 04 | Anthropic: Claude Fable 5.1Anthropic | 78.1 | 93.7 | 43.8 | 91.1 | 91.2 | 71 |
| 05 | 76.7 | 92.6 | 42.6 | 89.7 | 90.5 | 68 | |
| 06 | Llama 4 405BMeta | 76.4 | 71.2 | — | — | 86.9 | — |
| 07 | Claude Opus 5Anthropic | 75.8 | 91.8 | 40.6 | 89.0 | 90.6 | 67 |
| 08 | Qwen: Qwen3.8 Max (0902)Alibaba | 75.7 | 91.5 | 39.7 | 89.4 | 90.8 | 68 |
| 09 | OpenAI: GPT-5.6 SolOpenAI | 75.3 | 90.6 | 41.4 | 88.6 | 89.8 | 66 |
| 10 | OpenAI: GPT-5.5 ProOpenAI | 74.6 | 90.4 | 40.1 | 87.8 | 89.6 | 65 |
| 11 | Anthropic: Claude Fable 5Anthropic | 74.5 | 90.1 | 39.5 | 88.1 | 89.7 | 65 |
| 12 | MoonshotAI: Kimi K3Moonshot | 74.4 | 90.8 | 37.2 | 88.5 | 89.9 | 67 |
| 13 | Qwen 3 MaxAlibaba | 74.4 | 68.9 | — | — | 85.4 | — |
| 14 | 73.5 | 89.8 | 37.9 | 86.9 | 88.8 | 64 | |
| 15 | 73.4 | 89.9 | 36.1 | 87.6 | 89.3 | 65 | |
| 16 | Z.ai: GLM 5.3Z.ai | 73.2 | 89.7 | 34.8 | 87.9 | 89.1 | 66 |
| 17 | Anthropic: Claude Opus 4.8Anthropic | 72.8 | 89.1 | 36.3 | 86.5 | 88.5 | 64 |
| 18 | Mistral Large 3Mistral | 72.5 | 66.4 | — | — | 84.7 | — |
| 19 | Sakana: Fugu UltraSakana | 72.4 | 85.9 | 38.6 | 85.8 | 87.7 | 67 |
| 20 | DeepSeek: DeepSeek V4 Pro 0813DeepSeek | 72.3 | 88.8 | 33.5 | 87.1 | 88.9 | 65 |
| 21 | 72.3 | 89.1 | 35.8 | 83.0 | 89.2 | 66 | |
| 22 | Qwen: Qwen3.8 2.4T A95BAlibaba | 71.8 | 88.4 | 32.8 | 86.7 | 88.7 | 64 |
| 23 | OpenAI: GPT-5.5OpenAI | 71.5 | 88.1 | 34.6 | 85.4 | 87.9 | 62 |
| 24 | MoonshotAI: Kimi K2.6Moonshot | 71.2 | 88.2 | 32.6 | 85.9 | 87.5 | 63 |
| 25 | Google: Gemini 3.8 FlashGoogle | 71.0 | 85.9 | 39.8 | 80.0 | 86.8 | 63 |
| 26 | OpenAI: GPT-5.6 Sol ProOpenAI | 70.9 | 91.0 | 22.8 | 86.3 | 91.3 | 69 |
| 27 | Z.ai: GLM 5.2Z.ai | 70.6 | 87.6 | 31.4 | 85.6 | 87.8 | 62 |
| 28 | 70.3 | 83.9 | 35.6 | 83.8 | 86.5 | 65 | |
| 29 | Anthropic: Claude Sonnet 5Anthropic | 70.3 | 85.5 | 38.7 | 79.0 | 86.2 | 62 |
| 30 | DeepSeek: DeepSeek V4 Pro 0423DeepSeek | 69.3 | 86.2 | 29.8 | 84.5 | 86.8 | 61 |
| 31 | Qwen: Qwen3.7 MaxAlibaba | 69.3 | 86.5 | 29.7 | 84.2 | 87.1 | 60 |
| 32 | 68.9 | 85.8 | 30.4 | 83.6 | 86.1 | 59 | |
| 33 | OpenAI: GPT-5.4OpenAI | 68.8 | 88.8 | 22.7 | 80.7 | 88.2 | 71 |
| 34 | Z.ai: GLM 5Z.ai | 67.6 | 84.7 | 27.9 | 82.7 | 85.6 | 58 |
| 35 | Qwen: Qwen3 Max ThinkingAlibaba | 66.8 | 84.1 | 26.8 | 81.9 | 85.1 | 57 |
| 36 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 63.0 | 86.7 | 14.3 | 78.8 | 85.2 | 49 |
| 37 | 62.9 | 85.1 | 12.4 | 76.8 | 84.8 | 62 | |
| 38 | 62.7 | 81.9 | 19.8 | 73.5 | 81.6 | 62 | |
| 39 | Tencent: Hy3Tencent | 62.6 | 80.7 | 20.6 | 73.8 | 82.9 | 60 |
| 40 | Mistral: Mistral Medium 3.5Mistral | 61.8 | 81.2 | 18.2 | 73.8 | 83.6 | 56 |
| 41 | DeepSeek: DeepSeek V4 Flash 0731DeepSeek | 61.7 | 80.0 | 17.0 | 75.0 | 83.0 | 60 |
| 42 | 61.5 | 77.9 | 26.9 | 71.1 | 77.8 | 55 | |
| 43 | Qwen: Qwen3.6 Max PreviewAlibaba | 61.1 | 77.0 | 19.0 | 71.0 | 79.0 | 72 |
| 44 | 61.1 | 80.5 | 15.2 | 72.8 | 81.9 | 63 | |
| 45 | MiniMax: MiniMax M3MiniMax | 59.7 | 80.8 | 12.1 | 71.3 | 81.6 | 59 |
| 46 | Qwen: Qwen3.5 397B A17BAlibaba | 59.4 | 81.2 | 11.5 | 72.4 | 82.5 | 53 |
| 47 | 57.4 | 76.8 | 13.1 | 71.4 | 76.6 | 53 | |
| 48 | DeepSeek: DeepSeek V4 FlashDeepSeek | 51.6 | 70.0 | 7.0 | 62.0 | 75.0 | 52 |
How to read the reasoning leaderboard
Reasoning is where the gap between benchmark score and lived experience is widest, because most reasoning evaluations are multiple choice and most real reasoning is not. A model that scores 88% on GPQA Diamond has demonstrated that it can select the right graduate-level answer from ten options; it has not demonstrated that it will notice a flawed premise in your brief, or say 'this cannot be determined from the data given'.
GPQA Diamond still earns the heaviest weight. Its 448 questions were written by PhD holders and filtered so that skilled non-experts with full web access score around a third, which makes it genuinely resistant to lookup and shallow pattern matching. Its weaknesses are structural: multiple choice permits elimination strategies, the item count means a few points is noise, and its popularity guarantees creeping contamination.
ARC-AGI-2 is the counterweight. Its puzzles follow rules the model has never seen, constructed so memorisation cannot help, which makes it the most contamination-proof signal available and the best window onto genuine generalisation. It is also the least representative of production work — coloured grids resemble nothing you will ship — and high scores often come from enormous per-task compute budgets that are economically irrelevant. We weight it meaningfully but never decisively.
LiveBench supplies rotating, uncontaminated breadth, MMLU-Pro provides a wide knowledge floor, and the Artificial Analysis Intelligence Index adds an independently-executed cross-check run against production API endpoints rather than lab-side numbers. When a lab's self-reported figure and an independently-run one disagree sharply, we trust the independent one.
For picking a model, the honest advice is to test hard reasoning on your own material and watch behaviour rather than accuracy: does it show its work, does it flag uncertainty, does it refuse to invent a citation, and does it stay coherent when the problem is under-specified? None of that is on the leaderboard, and all of it decides whether the model is usable.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- What is the SOTA reasoning LLM?
- The current state-of-the-art reasoning model on our composite is displayed at the top of the table. Rankings are aggregated across GPQA, ARC-AGI, LiveBench, MMLU-Pro and Artificial Analysis.
- What does 'SOTA' mean in AI?
- SOTA stands for 'state of the art' — the model that currently leads on the benchmarks a given task cares about. For reasoning, it's the model highest on GPQA Diamond, ARC-AGI-2 and LiveBench.