Coding leaderboard
Best LLMs for coding
This leaderboard weights the benchmarks that actually predict real developer productivity: SWE-bench Verified (real GitHub bugs), Aider Polyglot (multi-language edits), Terminal-Bench (end-to-end shell agents), and LiveBench (contamination-resistant reasoning). Scores update every refresh cycle.
Composite weightingHow this is calculated →
Category leader
#01Anthropic
Anthropic: Claude Fable 5.
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Composite score
79.1
/ 100
+1.7 vs #2
Ranking
N=20| # | Model | Composite | SWE-bench | Aider Polyglot | Terminal-Bench | LiveBench | AA Index |
|---|---|---|---|---|---|---|---|
| 01 | Anthropic: Claude Fable 5Anthropic | 79.1 | 82.1 | 88.4 | 59.8 | 84.6 | 72 |
| 02 | MoonshotAI: Kimi K3Moonshot | 77.4 | 80.3 | 86.9 | 57.2 | 83.8 | 71 |
| 03 | 76.1 | 78.8 | 85.6 | 56.1 | 82.9 | 70 | |
| 04 | DeepSeek: DeepSeek V4 ProDeepSeek | 75.1 | 77.4 | 84.8 | 55.5 | 81.9 | 69 |
| 05 | Qwen: Qwen3.7 MaxAlibaba | 73.8 | 75.8 | 83.7 | 53.9 | 80.8 | 68 |
| 06 | Anthropic: Claude Opus 4.8Anthropic | 72.8 | 74.9 | 82.8 | 52.7 | 80.1 | 67 |
| 07 | Anthropic: Claude Opus 4.8 (Fast)Anthropic | 72.8 | 74.9 | 82.8 | 52.7 | 80.1 | 67 |
| 08 | OpenAI: GPT-5.6 SolOpenAI | 72.6 | 72.4 | 78.1 | 58.2 | 82.4 | 74 |
| 09 | Z.ai: GLM 5.2Z.ai | 71.8 | 73.6 | 81.9 | 51.8 | 79.4 | 66 |
| 10 | NVIDIA: Nemotron 3 UltraNVIDIA | 70.6 | 72.2 | 80.8 | 50.6 | 78.6 | 65 |
| 11 | Sakana: Fugu UltraSakana | 69.6 | 70.8 | 79.9 | 49.7 | 77.9 | 64 |
| 12 | 69.1 | 70.4 | 79.4 | 49.3 | 77.5 | 63 | |
| 13 | OpenAI: GPT-5.6 TerraOpenAI | 68.8 | 69.9 | 79.1 | 48.9 | 77.2 | 63 |
| 14 | MoonshotAI: Kimi K2.7 CodeMoonshot | 68.2 | 70.1 | 79.6 | 47.4 | 75.3 | 60 |
| 15 | 67.9 | 68.8 | 78.6 | 47.9 | 76.5 | 62 | |
| 16 | MiniMax: MiniMax M3MiniMax | 66.9 | 67.6 | 77.8 | 46.8 | 75.8 | 61 |
| 17 | Z.ai: GLM 5.1Z.ai | 66.0 | 66.9 | 77.1 | 45.9 | 74.5 | 59 |
| 18 | 65.2 | 65.8 | 76.4 | 45.2 | 73.9 | 58 | |
| 19 | MoonshotAI: Kimi K2.6Moonshot | 64.5 | 65.2 | 75.9 | 44.6 | 73.2 | 57 |
| 20 | Mistral: Mistral Medium 3.5Mistral | 63.3 | 63.7 | 74.8 | 43.4 | 72.3 | 56 |
Frequently asked
- What is the best LLM for coding right now?
- The current #1 on this coding-weighted composite is shown at the top of the table. It aggregates SWE-bench Verified (35%), Aider Polyglot (25%), Terminal-Bench (20%), LiveBench (15%) and AA Index (5%).
- Why isn't a benchmark like HumanEval used?
- HumanEval is saturated and heavily contaminated in modern training data, so it no longer discriminates between frontier models. SWE-bench Verified and Aider Polyglot are today's stronger signals.