Anthropic: Claude Fable 5.1
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Best value in each row highlighted in accent. Blank cells mean the provider hasn't reported that benchmark yet.
| Benchmark | Anthropic: Claude Fable 5.1 | Qwen: Qwen3.8 Max (0902) | SpaceXAI: Grok 4.6 | MoonshotAI: Kimi K3 | Z.ai: GLM 5.3 |
|---|---|---|---|---|---|
Arena Elo Human blind-vote Elo from LMArena — the closest proxy to 'which model feels smarter'. | 1638 | 1619 | 1627 | 1608 | 1604 |
AA Index Artificial Analysis' composite Intelligence Index across reasoning, math and coding. | 71.0 | 68.0 | 68.0 | 67.0 | 66.0 |
LiveBench Contamination-resistant benchmark refreshed monthly — a strong signal of raw capability. | 91.1 | 89.4 | 89.7 | 88.5 | 87.9 |
GPQA Graduate-level science QA. Above 70% is frontier reasoning territory. | 93.7% | 91.5% | 92.6% | 90.8% | 89.7% |
MMLU-Pro Broad knowledge across 57 subjects — the classic capability floor. | 91.2% | 90.8% | 90.5% | 89.9% | 89.1% |
SWE-Bench Real-world GitHub bug-fixing. The single best proxy for agentic coding. | 86.1% | 78.7% | 80.3% | 79.4% | 77.8% |
Terminal-Bench End-to-end shell tasks. Measures tool-use and long-horizon execution. | 77.4% | 69.3% | 71.8% | 70.2% | 70.6% |
ARC-AGI Abstract reasoning puzzles. Historically brutal — even small gains matter. | 43.8% | 39.7% | 42.6% | 37.2% | 34.8% |
Aider Polyglot Multi-language code editing benchmark. Signals practical dev-loop quality. | 89.6% | 84.9% | 85.1% | 86.8% | 86.3% |
Context Maximum input tokens. Bigger unlocks whole-codebase and long-doc workflows. | 1.0M | 1.0M | 500K | 1.0M | 1.3M |
Speed (tok/s) Output tokens per second. Matters for interactive UIs and long generations. | 63 | 73 | 116 | 58 | 71 |
Input $/1M Cost per million input tokens. | $10.00 | $2.00 | $2.00 | $3.00 | $1.40 |
Output $/1M Cost per million output tokens. | $50.00 | $6.00 | $6.00 | $15.00 | $4.40 |
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Optimize for SWE-Bench and Terminal-Bench. Aider Polyglot is the tiebreaker for real dev-loop quality.
GPQA and ARC-AGI matter more than MMLU. A high Arena Elo helps for open-ended prompts.
Filter by context window first (200K+ for full-repo work), then compare LiveBench to avoid quality drop-off.
Input and output $/1M dominate at scale. Sort by price, then take the highest AA Index still in budget.