Anthropic: Claude Fable 5
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Best value in each row highlighted in accent. Blank cells mean the provider hasn't reported that benchmark yet.
| Benchmark | Anthropic: Claude Fable 5 | MoonshotAI: Kimi K3 | OpenAI: GPT-5.6 Sol | xAI: Grok 4.5 | DeepSeek: DeepSeek V4 Pro |
|---|---|---|---|---|---|
Arena Elo Human blind-vote Elo from LMArena — the closest proxy to 'which model feels smarter'. | 1572 | 1558 | 1560 | 1564 | 1539 |
AA Index Artificial Analysis' composite Intelligence Index across reasoning, math and coding. | 72.0 | 71.0 | 74.0 | 70.0 | 69.0 |
LiveBench Contamination-resistant benchmark refreshed monthly — a strong signal of raw capability. | 84.6 | 83.8 | 82.4 | 82.9 | 81.9 |
GPQA Graduate-level science QA. Above 70% is frontier reasoning territory. | 91.2% | 90.4% | 84.1% | 89.7% | 88.9% |
MMLU-Pro Broad knowledge across 57 subjects — the classic capability floor. | 91.4% | 90.9% | 91.5% | 90.1% | 89.6% |
SWE-Bench Real-world GitHub bug-fixing. The single best proxy for agentic coding. | 82.1% | 80.3% | 72.4% | 78.8% | 77.4% |
Terminal-Bench End-to-end shell tasks. Measures tool-use and long-horizon execution. | 59.8% | 57.2% | 58.2% | 56.1% | 55.5% |
ARC-AGI Abstract reasoning puzzles. Historically brutal — even small gains matter. | 23.8% | 22.7% | 46.0% | 22.1% | 20.8% |
Aider Polyglot Multi-language code editing benchmark. Signals practical dev-loop quality. | 88.4% | 86.9% | 78.1% | 85.6% | 84.8% |
Context Maximum input tokens. Bigger unlocks whole-codebase and long-doc workflows. | 1.0M | 1.0M | 1.1M | 500K | 1.0M |
Speed (tok/s) Output tokens per second. Matters for interactive UIs and long generations. | 42 | 55 | 92 | 76 | 91 |
Input $/1M Cost per million input tokens. | $10.00 | $3.00 | $5.00 | $2.00 | $0.43 |
Output $/1M Cost per million output tokens. | $50.00 | $15.00 | $30.00 | $6.00 | $0.87 |
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
Optimize for SWE-Bench and Terminal-Bench. Aider Polyglot is the tiebreaker for real dev-loop quality.
GPQA and ARC-AGI matter more than MMLU. A high Arena Elo helps for open-ended prompts.
Filter by context window first (200K+ for full-repo work), then compare LiveBench to avoid quality drop-off.
Input and output $/1M dominate at scale. Sort by price, then take the highest AA Index still in budget.