Reasoning leaderboard
Best LLMs for reasoning
Reasoning-heavy tasks — graduate science QA, abstract puzzles, novel logic — separate frontier models from the pack. This ranking weights GPQA Diamond (30%), ARC-AGI-2 (25%), LiveBench (20%), MMLU-Pro (15%) and AA Index (10%).
Composite weightingHow this is calculated →
Category leader
#01OpenAI
OpenAI: GPT-5.6 Sol.
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Composite score
74.3
/ 100
+3.2 vs #2
Ranking
N=20| # | Model | Composite | GPQA | ARC-AGI | LiveBench | MMLU-Pro | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | OpenAI: GPT-5.6 SolOpenAI | 74.3 | 84.1 | 46.0 | 82.4 | 91.5 | 74 |
| 02 | Anthropic: Claude Fable 5Anthropic | 71.1 | 91.2 | 23.8 | 84.6 | 91.4 | 72 |
| 03 | MoonshotAI: Kimi K3Moonshot | 70.3 | 90.4 | 22.7 | 83.8 | 90.9 | 71 |
| 04 | 69.5 | 89.7 | 22.1 | 82.9 | 90.1 | 70 | |
| 05 | DeepSeek: DeepSeek V4 ProDeepSeek | 68.6 | 88.9 | 20.8 | 81.9 | 89.6 | 69 |
| 06 | Qwen: Qwen3.7 MaxAlibaba | 67.8 | 88.1 | 20.2 | 80.8 | 88.9 | 68 |
| 07 | Anthropic: Claude Opus 4.8Anthropic | 67.2 | 87.6 | 19.6 | 80.1 | 88.4 | 67 |
| 08 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 67.2 | 87.6 | 19.6 | 80.1 | 88.4 | 67 |
| 09 | Z.ai: GLM 5.2Z.ai | 66.4 | 86.9 | 18.9 | 79.4 | 87.8 | 66 |
| 10 | NVIDIA: Nemotron 3 UltraNVIDIA | 65.7 | 86.1 | 18.2 | 78.6 | 87.2 | 65 |
| 11 | Sakana: Fugu UltraSakana | 65.0 | 85.4 | 17.7 | 77.9 | 86.7 | 64 |
| 12 | 64.6 | 85.1 | 17.3 | 77.5 | 86.4 | 63 | |
| 13 | OpenAI: GPT-5.6 TerraOpenAI | 64.4 | 84.8 | 17.1 | 77.2 | 86.1 | 63 |
| 14 | 63.8 | 84.2 | 16.8 | 76.5 | 85.6 | 62 | |
| 15 | MiniMax: MiniMax M3MiniMax | 63.2 | 83.6 | 16.2 | 75.8 | 85.1 | 61 |
| 16 | MoonshotAI: Kimi K2.7 CodeMoonshot | 62.6 | 82.9 | 15.8 | 75.3 | 84.7 | 60 |
| 17 | Z.ai: GLM 5.1Z.ai | 61.9 | 82.1 | 15.3 | 74.5 | 84.0 | 59 |
| 18 | 61.4 | 81.7 | 14.9 | 73.9 | 83.6 | 58 | |
| 19 | MoonshotAI: Kimi K2.6Moonshot | 60.7 | 80.9 | 14.5 | 73.2 | 83.1 | 57 |
| 20 | Mistral: Mistral Medium 3.5Mistral | 59.9 | 80.1 | 13.8 | 72.3 | 82.5 | 56 |
Frequently asked
- What is the SOTA reasoning LLM?
- The current state-of-the-art reasoning model on our composite is displayed at the top of the table. Rankings are aggregated across GPQA, ARC-AGI, LiveBench, MMLU-Pro and Artificial Analysis.
- What does 'SOTA' mean in AI?
- SOTA stands for 'state of the art' — the model that currently leads on the benchmarks a given task cares about. For reasoning, it's the model highest on GPQA Diamond, ARC-AGI-2 and LiveBench.