Agentic leaderboard
Best agentic LLMs
Agentic workloads reward long-horizon planning, tool-use and recovery from errors. This ranking weights Terminal-Bench (40%), SWE-bench Verified (30%), Aider Polyglot (15%), LiveBench (10%) and AA Index (5%).
Category leader
Anthropic: Claude Fable 5.1.
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
+4.0 vs #2
Ranking
N=42| # | Model | Composite | Terminal-Bench | SWE-bench | Aider Polyglot | LiveBench | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | Anthropic: Claude Fable 5.1Anthropic | 82.9 | 77.4 | 86.1 | 89.6 | 91.1 | 71 |
| 02 | Claude Opus 5Anthropic | 78.9 | 72.7 | 81.6 | 87.1 | 89.0 | 67 |
| 03 | 77.9 | 71.8 | 80.3 | 85.1 | 89.7 | 68 | |
| 04 | Anthropic: Claude Fable 5Anthropic | 77.7 | 71.6 | 80.2 | 86.2 | 88.1 | 65 |
| 05 | OpenAI: GPT-5.6 SolOpenAI | 77.3 | 72.1 | 78.8 | 84.5 | 88.6 | 66 |
| 06 | MoonshotAI: Kimi K3Moonshot | 77.1 | 70.2 | 79.4 | 86.8 | 88.5 | 67 |
| 07 | Z.ai: GLM 5.3Z.ai | 76.6 | 70.6 | 77.8 | 86.3 | 87.9 | 66 |
| 08 | Qwen: Qwen3.8 Max (0902)Alibaba | 76.4 | 69.3 | 78.7 | 84.9 | 89.4 | 68 |
| 09 | OpenAI: GPT-5.5 ProOpenAI | 75.8 | 70.3 | 77.4 | 82.8 | 87.8 | 65 |
| 10 | Anthropic: Claude Opus 4.8Anthropic | 75.8 | 69.1 | 78.6 | 84.8 | 86.5 | 64 |
| 11 | OpenAI: GPT-5.6 Sol ProOpenAI | 75.6 | 68.5 | 80.3 | 80.1 | 86.3 | 69 |
| 12 | 75.2 | 68.9 | 76.9 | 83.7 | 87.6 | 65 | |
| 13 | DeepSeek: DeepSeek V4 Pro 0813DeepSeek | 74.7 | 67.5 | 76.4 | 85.6 | 87.1 | 65 |
| 14 | 74.4 | 67.7 | 76.8 | 82.6 | 86.9 | 64 | |
| 15 | MoonshotAI: Kimi K2.6Moonshot | 73.5 | 65.9 | 75.8 | 84.1 | 85.9 | 63 |
| 16 | Qwen: Qwen3.8 2.4T A95BAlibaba | 72.8 | 65.2 | 74.8 | 82.4 | 86.7 | 64 |
| 17 | 72.7 | 63.9 | 76.4 | 83.8 | 83.0 | 66 | |
| 18 | OpenAI: GPT-5.5OpenAI | 72.4 | 66.2 | 74.1 | 80.1 | 85.4 | 62 |
| 19 | Z.ai: GLM 5.2Z.ai | 72.1 | 64.8 | 73.9 | 82.2 | 85.6 | 62 |
| 20 | Anthropic: Claude Sonnet 5Anthropic | 72.0 | 62.7 | 77.6 | 84.0 | 79.0 | 62 |
| 21 | Sakana: Fugu UltraSakana | 71.9 | 63.7 | 74.6 | 80.8 | 85.8 | 67 |
| 22 | Google: Gemini 3.8 FlashGoogle | 70.6 | 61.5 | 76.1 | 80.4 | 80.0 | 63 |
| 23 | DeepSeek: DeepSeek V4 Pro 0423DeepSeek | 70.4 | 62.7 | 72.6 | 80.6 | 84.5 | 61 |
| 24 | 70.3 | 62.6 | 74.7 | 80.8 | 76.8 | 62 | |
| 25 | OpenAI: GPT-5.4OpenAI | 69.8 | 61.0 | 73.7 | 77.8 | 80.7 | 71 |
| 26 | Qwen: Qwen3.7 MaxAlibaba | 69.4 | 61.8 | 71.2 | 79.5 | 84.2 | 60 |
| 27 | 69.2 | 60.3 | 71.8 | 79.2 | 83.8 | 65 | |
| 28 | 68.8 | 61.4 | 70.8 | 77.8 | 83.6 | 59 | |
| 29 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 68.5 | 60.9 | 72.8 | 79.8 | 78.8 | 49 |
| 30 | Z.ai: GLM 5Z.ai | 68.0 | 60.3 | 69.7 | 78.9 | 82.7 | 58 |
| 31 | Qwen: Qwen3 Max ThinkingAlibaba | 66.6 | 58.7 | 68.4 | 77.4 | 81.9 | 57 |
| 32 | DeepSeek: DeepSeek V4 Flash 0731DeepSeek | 66.4 | 57.0 | 70.0 | 81.0 | 75.0 | 60 |
| 33 | 62.1 | 52.4 | 63.5 | 77.4 | 73.5 | 62 | |
| 34 | MiniMax: MiniMax M3MiniMax | 61.7 | 53.2 | 65.1 | 71.9 | 71.3 | 59 |
| 35 | Qwen: Qwen3.5 397B A17BAlibaba | 60.9 | 50.8 | 64.7 | 75.4 | 72.4 | 53 |
| 36 | 60.8 | 49.4 | 64.6 | 74.6 | 72.8 | 63 | |
| 37 | Mistral: Mistral Medium 3.5Mistral | 60.4 | 52.8 | 64.3 | 65.4 | 73.8 | 56 |
| 38 | 59.9 | 48.5 | 64.1 | 76.1 | 71.1 | 55 | |
| 39 | Qwen: Qwen3.6 Max PreviewAlibaba | 59.8 | 50.0 | 60.0 | 74.0 | 71.0 | 72 |
| 40 | Tencent: Hy3Tencent | 58.4 | 47.2 | 62.1 | 70.2 | 73.8 | 60 |
| 41 | 52.2 | 38.1 | 54.7 | 71.5 | 71.4 | 53 | |
| 42 | DeepSeek: DeepSeek V4 FlashDeepSeek | 50.4 | 38.0 | 53.0 | 70.0 | 62.0 | 52 |
How to read the agentic leaderboard
An agent is a model that has to survive its own mistakes. Chat quality barely predicts this. What matters is whether the model can hold a plan across dozens of steps, choose the right tool, read an error message honestly, revise its approach, and stop when the job is done rather than looping until the budget runs out.
Terminal-Bench dominates the weighting because it is the only widely-used evaluation that puts the model in a real shell and grades the end state of the environment. Absolute scores look low compared with other benchmarks, and that is the point: unattended multi-step work is genuinely hard, and a benchmark where everyone scores 90% would tell you nothing. SWE-bench Verified adds a second execution-graded signal with an unforgeable pass criterion, and Aider Polyglot checks that the model's edits are well-formed enough for a tool loop to apply without human repair.
Two failure modes are invisible on this table and worth testing yourself. The first is tool-call reliability — malformed arguments, hallucinated function names, or silently skipping a required call — which sinks agents far more often than weak reasoning does. The second is knowing when to stop: models that cannot recognise completion burn tokens indefinitely and models that give up early leave tasks half-finished, and neither behaviour shows up as a benchmark delta.
Cost behaves differently for agents too. A single agentic task can consume hundreds of thousands of tokens across retries and hidden reasoning, so a model with a higher per-token price but a higher first-pass success rate is frequently cheaper per completed task. Measure completions, not tokens, and pair this leaderboard with the price and throughput tables before you commit to a production model.
Finally, harness matters as much as the model. The same model can move ten points or more depending on retry budget, context management and tool design, so use this ranking to shortlist candidates and then benchmark them inside your own agent loop.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- What is the best agentic AI model?
- The top of this table shows the current leader on our agentic composite. Terminal-Bench and SWE-bench Verified dominate the weighting because they measure real multi-step task completion, not one-shot answers.
- Are these live scores?
- Yes — the underlying benchmark data is refreshed every couple of hours from the public sources listed at the bottom of the page.