Agentic leaderboard

Best agentic LLMs

Agentic workloads reward long-horizon planning, tool-use and recovery from errors. This ranking weights Terminal-Bench (40%), SWE-bench Verified (30%), Aider Polyglot (15%), LiveBench (10%) and AA Index (5%).

Composite weightingHow this is calculated →

Category leader

#01Anthropic

Anthropic: Claude Fable 5.

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Composite score
73.9
/ 100

+1.9 vs #2

Ranking

N=20
#ModelCompositeTerminal-BenchSWE-benchAider PolyglotLiveBenchAA Index
0173.959.882.188.484.672
0271.957.280.386.983.871
0370.756.178.885.682.970
0469.855.577.484.881.969
0568.758.272.478.182.474
0668.353.975.883.780.868
0767.352.774.982.880.167
0867.352.774.982.880.167
0966.351.873.681.979.466
1065.150.672.280.878.665
1164.149.770.879.977.964
1263.649.370.479.477.563
1363.348.969.979.177.263
1462.547.470.179.675.360
1562.347.968.878.676.562
1661.346.867.677.875.861
1760.445.966.977.174.559
1859.645.265.876.473.958
1959.044.665.275.973.257
2057.743.463.774.872.356

Frequently asked

What is the best agentic AI model?
The top of this table shows the current leader on our agentic composite. Terminal-Bench and SWE-bench Verified dominate the weighting because they measure real multi-step task completion, not one-shot answers.
Are these live scores?
Yes — the underlying benchmark data is refreshed every couple of hours from the public sources listed at the bottom of the page.