Agentic leaderboard
Best agentic LLMs
Agentic workloads reward long-horizon planning, tool-use and recovery from errors. This ranking weights Terminal-Bench (40%), SWE-bench Verified (30%), Aider Polyglot (15%), LiveBench (10%) and AA Index (5%).
Composite weightingHow this is calculated →
Category leader
#01Anthropic
Anthropic: Claude Fable 5.
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Composite score
73.9
/ 100
+1.9 vs #2
Ranking
N=20| # | Model | Composite | Terminal-Bench | SWE-bench | Aider Polyglot | LiveBench | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | Anthropic: Claude Fable 5Anthropic | 73.9 | 59.8 | 82.1 | 88.4 | 84.6 | 72 |
| 02 | MoonshotAI: Kimi K3Moonshot | 71.9 | 57.2 | 80.3 | 86.9 | 83.8 | 71 |
| 03 | 70.7 | 56.1 | 78.8 | 85.6 | 82.9 | 70 | |
| 04 | DeepSeek: DeepSeek V4 ProDeepSeek | 69.8 | 55.5 | 77.4 | 84.8 | 81.9 | 69 |
| 05 | OpenAI: GPT-5.6 SolOpenAI | 68.7 | 58.2 | 72.4 | 78.1 | 82.4 | 74 |
| 06 | Qwen: Qwen3.7 MaxAlibaba | 68.3 | 53.9 | 75.8 | 83.7 | 80.8 | 68 |
| 07 | Anthropic: Claude Opus 4.8Anthropic | 67.3 | 52.7 | 74.9 | 82.8 | 80.1 | 67 |
| 08 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 67.3 | 52.7 | 74.9 | 82.8 | 80.1 | 67 |
| 09 | Z.ai: GLM 5.2Z.ai | 66.3 | 51.8 | 73.6 | 81.9 | 79.4 | 66 |
| 10 | NVIDIA: Nemotron 3 UltraNVIDIA | 65.1 | 50.6 | 72.2 | 80.8 | 78.6 | 65 |
| 11 | Sakana: Fugu UltraSakana | 64.1 | 49.7 | 70.8 | 79.9 | 77.9 | 64 |
| 12 | 63.6 | 49.3 | 70.4 | 79.4 | 77.5 | 63 | |
| 13 | OpenAI: GPT-5.6 TerraOpenAI | 63.3 | 48.9 | 69.9 | 79.1 | 77.2 | 63 |
| 14 | MoonshotAI: Kimi K2.7 CodeMoonshot | 62.5 | 47.4 | 70.1 | 79.6 | 75.3 | 60 |
| 15 | 62.3 | 47.9 | 68.8 | 78.6 | 76.5 | 62 | |
| 16 | MiniMax: MiniMax M3MiniMax | 61.3 | 46.8 | 67.6 | 77.8 | 75.8 | 61 |
| 17 | Z.ai: GLM 5.1Z.ai | 60.4 | 45.9 | 66.9 | 77.1 | 74.5 | 59 |
| 18 | 59.6 | 45.2 | 65.8 | 76.4 | 73.9 | 58 | |
| 19 | MoonshotAI: Kimi K2.6Moonshot | 59.0 | 44.6 | 65.2 | 75.9 | 73.2 | 57 |
| 20 | Mistral: Mistral Medium 3.5Mistral | 57.7 | 43.4 | 63.7 | 74.8 | 72.3 | 56 |
Frequently asked
- What is the best agentic AI model?
- The top of this table shows the current leader on our agentic composite. Terminal-Bench and SWE-bench Verified dominate the weighting because they measure real multi-step task completion, not one-shot answers.
- Are these live scores?
- Yes — the underlying benchmark data is refreshed every couple of hours from the public sources listed at the bottom of the page.