Reasoning leaderboard

Best LLMs for reasoning

Reasoning-heavy tasks — graduate science QA, abstract puzzles, novel logic — separate frontier models from the pack. This ranking weights GPQA Diamond (30%), ARC-AGI-2 (25%), LiveBench (20%), MMLU-Pro (15%) and AA Index (10%).

Composite weightingHow this is calculated →

Category leader

#01OpenAI

OpenAI: GPT-5.6 Sol.

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Composite score
74.3
/ 100

+3.2 vs #2

Ranking

N=20
#ModelCompositeGPQAARC-AGILiveBenchMMLU-ProAA Index
0174.384.146.082.491.574
0271.191.223.884.691.472
0370.390.422.783.890.971
0469.589.722.182.990.170
0568.688.920.881.989.669
0667.888.120.280.888.968
0767.287.619.680.188.467
0867.287.619.680.188.467
0966.486.918.979.487.866
1065.786.118.278.687.265
1165.085.417.777.986.764
1264.685.117.377.586.463
1364.484.817.177.286.163
1463.884.216.876.585.662
1563.283.616.275.885.161
1662.682.915.875.384.760
1761.982.115.374.584.059
1861.481.714.973.983.658
1960.780.914.573.283.157
2059.980.113.872.382.556

Frequently asked

What is the SOTA reasoning LLM?
The current state-of-the-art reasoning model on our composite is displayed at the top of the table. Rankings are aggregated across GPQA, ARC-AGI, LiveBench, MMLU-Pro and Artificial Analysis.
What does 'SOTA' mean in AI?
SOTA stands for 'state of the art' — the model that currently leads on the benchmarks a given task cares about. For reasoning, it's the model highest on GPQA Diamond, ARC-AGI-2 and LiveBench.