Coding leaderboard

Best LLMs for coding

This leaderboard weights the benchmarks that actually predict real developer productivity: SWE-bench Verified (real GitHub bugs), Aider Polyglot (multi-language edits), Terminal-Bench (end-to-end shell agents), and LiveBench (contamination-resistant reasoning). Scores update every refresh cycle.

Composite weightingHow this is calculated →

Category leader

#01Anthropic

Anthropic: Claude Fable 5.

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Composite score
79.1
/ 100

+1.7 vs #2

Ranking

N=20
#ModelCompositeSWE-benchAider PolyglotTerminal-BenchLiveBenchAA Index
0179.182.188.459.884.672
0277.480.386.957.283.871
0376.178.885.656.182.970
0475.177.484.855.581.969
0573.875.883.753.980.868
0672.874.982.852.780.167
0772.874.982.852.780.167
0872.672.478.158.282.474
0971.873.681.951.879.466
1070.672.280.850.678.665
1169.670.879.949.777.964
1269.170.479.449.377.563
1368.869.979.148.977.263
1468.270.179.647.475.360
1567.968.878.647.976.562
1666.967.677.846.875.861
1766.066.977.145.974.559
1865.265.876.445.273.958
1964.565.275.944.673.257
2063.363.774.843.472.356

Frequently asked

What is the best LLM for coding right now?
The current #1 on this coding-weighted composite is shown at the top of the table. It aggregates SWE-bench Verified (35%), Aider Polyglot (25%), Terminal-Bench (20%), LiveBench (15%) and AA Index (5%).
Why isn't a benchmark like HumanEval used?
HumanEval is saturated and heavily contaminated in modern training data, so it no longer discriminates between frontier models. SWE-bench Verified and Aider Polyglot are today's stronger signals.