Coding leaderboard
Best LLMs for coding
This leaderboard weights the benchmarks that actually predict real developer productivity: SWE-bench Verified (real GitHub bugs), Aider Polyglot (multi-language edits), Terminal-Bench (end-to-end shell agents), and LiveBench (contamination-resistant reasoning). Scores update every refresh cycle.
Category leader
Anthropic: Claude Fable 5.1.
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
+4.5 vs #2
Ranking
N=43| # | Model | Composite | SWE-bench | Aider Polyglot | Terminal-Bench | LiveBench | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | Anthropic: Claude Fable 5.1Anthropic | 84.7 | 86.8 | 88.1 | 74.2 | 90.8 | 77 |
| 02 | Qwen: Qwen3.8 Max (0902)Alibaba | 80.2 | 81.2 | 84.7 | 67.1 | 89.4 | 75 |
| 03 | 80.1 | 80.6 | 83.8 | 69.4 | 88.9 | 75 | |
| 04 | MoonshotAI: Kimi K3Moonshot | 79.8 | 79.8 | 85.2 | 68.6 | 87.9 | 73 |
| 05 | Z.ai: GLM 5.3Z.ai | 79.3 | 80.1 | 84.1 | 67.8 | 87.3 | 72 |
| 06 | Claude Opus 5Anthropic | 79.3 | 80.3 | 84.4 | 68.2 | 86.1 | 70 |
| 07 | DeepSeek: DeepSeek V4 Pro 0813DeepSeek | 78.0 | 78.4 | 83.4 | 65.7 | 86.8 | 71 |
| 08 | OpenAI: GPT-5.6 SolOpenAI | 77.8 | 78.1 | 82.9 | 66.2 | 86.5 | 71 |
| 09 | 77.5 | 77.9 | 82.7 | 65.4 | 86.3 | 70 | |
| 10 | Anthropic: Claude Sonnet 5Anthropic | 76.9 | 76.3 | 84.1 | 63.1 | 83.3 | 81 |
| 11 | OpenAI: GPT-5.5OpenAI | 75.7 | 76.4 | 80.2 | 63.8 | 85.1 | 68 |
| 12 | Google: Gemini 3.8 FlashGoogle | 75.4 | 74.1 | 81.9 | 61.6 | 84.1 | 82 |
| 13 | Qwen: Qwen3.7 MaxAlibaba | 74.3 | 74.1 | 80.4 | 61.7 | 83.8 | 67 |
| 14 | 73.5 | 74.7 | 80.8 | 62.6 | 76.8 | 62 | |
| 15 | OpenAI: GPT-5.4OpenAI | 73.1 | 73.7 | 77.8 | 61.0 | 80.7 | 71 |
| 16 | NVIDIA: Nemotron 3 UltraNVIDIA | 73.0 | 72.8 | 78.7 | 60.9 | 82.6 | 65 |
| 17 | Qwen: Qwen3.8 27BAlibaba | 72.4 | 71.8 | 78.6 | 57.9 | 81.2 | 78 |
| 18 | Anthropic: Claude Opus 4.8Anthropic | 72.3 | 72.4 | 79.2 | 60.1 | 80.0 | 62 |
| 19 | Sakana: Fugu UltraSakana | 72.3 | 71.9 | 77.9 | 60.4 | 82.2 | 64 |
| 20 | OpenAI: GPT-5.3-CodexOpenAI | 72.2 | 72.9 | 82.2 | 58.3 | 77.4 | 58 |
| 21 | 72.0 | 70.8 | 77.8 | 61.4 | 83.6 | 59 | |
| 22 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 71.9 | 72.8 | 79.8 | 60.9 | 78.8 | 49 |
| 23 | 71.7 | 70.9 | 77.8 | 57.2 | 80.7 | 78 | |
| 24 | MoonshotAI: Kimi K2.6Moonshot | 71.5 | 71.1 | 78.3 | 59.2 | 80.3 | 63 |
| 25 | 70.9 | 70.8 | 77.1 | 58.1 | 80.8 | 63 | |
| 26 | DeepSeek: DeepSeek V4 Pro 0423DeepSeek | 70.8 | 70.4 | 77.4 | 57.9 | 80.5 | 63 |
| 27 | 70.4 | 69.8 | 76.8 | 58.7 | 79.7 | 62 | |
| 28 | DeepSeek: DeepSeek V4 Flash 0731DeepSeek | 70.4 | 70.0 | 81.0 | 57.0 | 75.0 | 60 |
| 29 | Qwen: Qwen3 Max ThinkingAlibaba | 70.2 | 68.4 | 77.4 | 58.7 | 81.9 | 57 |
| 30 | Tencent: Hy4 previewTencent | 70.0 | 69.7 | 77.2 | 56.2 | 80.1 | 62 |
| 31 | Mistral: Devstral 2 2512Mistral | 70.0 | 70.1 | 78.1 | 57.1 | 76.9 | 59 |
| 32 | Z.ai: GLM 5.2Z.ai | 69.9 | 69.5 | 76.6 | 57.6 | 79.1 | 61 |
| 33 | Anthropic: Claude Opus 4.6Anthropic | 69.9 | 70.9 | 79.9 | 52.2 | 76.4 | 64 |
| 34 | 69.3 | 68.9 | 76.1 | 56.8 | 78.7 | 60 | |
| 35 | Qwen: Qwen3.8 2.4T A95BAlibaba | 68.8 | 68.3 | 75.8 | 56.1 | 78.3 | 60 |
| 36 | Thinking Machines: InklingThinking Machines | 68.2 | 67.4 | 75.4 | 55.7 | 77.8 | 59 |
| 37 | 66.2 | 63.5 | 77.4 | 52.4 | 73.5 | 62 | |
| 38 | Qwen: Qwen3.5 397B A17BAlibaba | 65.2 | 64.7 | 75.4 | 50.8 | 72.4 | 53 |
| 39 | MiniMax: MiniMax M3MiniMax | 65.0 | 65.1 | 71.9 | 53.2 | 71.3 | 59 |
| 40 | Mistral: Mistral Medium 3.5Mistral | 63.3 | 64.3 | 65.4 | 52.8 | 73.8 | 56 |
| 41 | Tencent: Hy3Tencent | 62.8 | 62.1 | 70.2 | 47.2 | 73.8 | 60 |
| 42 | 58.0 | 54.7 | 71.5 | 38.1 | 71.4 | 53 | |
| 43 | DeepSeek: DeepSeek V4 FlashDeepSeek | 55.5 | 53.0 | 70.0 | 38.0 | 62.0 | 52 |
How to read the coding leaderboard
Coding ability is the hardest capability to summarise in one number, because 'writing code' is really four separate skills: producing a correct patch, applying that patch to an existing file without mangling it, navigating an unfamiliar repository, and recovering when a test fails. Each of the benchmarks below measures a different one, which is why we weight rather than pick.
SWE-bench Verified carries the most weight because its pass criterion cannot be argued with: the model's patch either makes the project's own hidden tests go green or it does not. It is the closest public proxy for 'can this model close a real ticket'. Its limits are worth knowing — it is Python-only, drawn from a narrow set of large repositories, and results shift by ten points or more depending on the agent harness wrapped around the model. Compare scores from the same harness or not at all.
Aider Polyglot fills the language gap and, more importantly, measures edit-format compliance across C++, Go, Java, JavaScript, Python and Rust. A model that reasons brilliantly but emits malformed diffs will feel broken inside an IDE regardless of its SWE-bench number, and this is the benchmark that exposes that failure. Terminal-Bench then adds the long-horizon dimension: planning, running commands, reading stderr and course-correcting without a human in the loop. LiveBench acts as the contamination-resistant control, since its questions rotate and cannot have been memorised.
A practical way to use the table: shortlist the top three, then check price and throughput before committing. Coding workloads are token-hungry, and a model two points behind at a third of the cost is usually the better production choice. Reasoning models also emit large volumes of hidden thinking tokens, so measure cost per merged pull request rather than cost per token — the ordering often changes when you do.
Finally, treat gaps under roughly three points as ties. These evaluations have a few hundred items each, which puts several points inside the noise band. If two models are that close, the deciding factors should be latency, context window, tool-calling reliability and how they behave on your own repository — not the leaderboard.
More on how these numbers are produced in the methodology and what each evaluation measures in the benchmark guide. Spotted a score that disagrees with its source? Tell us.
Frequently asked
- What is the best LLM for coding right now?
- The current #1 on this coding-weighted composite is shown at the top of the table. It aggregates SWE-bench Verified (35%), Aider Polyglot (25%), Terminal-Bench (20%), LiveBench (15%) and AA Index (5%).
- Why isn't a benchmark like HumanEval used?
- HumanEval is saturated and heavily contaminated in modern training data, so it no longer discriminates between frontier models. SWE-bench Verified and Aider Polyglot are today's stronger signals.