System Analysis // Reference ID: 294-C

Anthropic: Claude Fable 5
vs the field

A forensic evaluation of today's LLM leaderboard. We pressure-test reasoning density, coding, and instruction adherence across five Tier-1 AI models — pulled live from the same aggregated feed the homepage uses.

The contenders

Primary Subject

Anthropic: Claude Fable 5

State of the ArtAnthropictextimagefile

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Contender 02 · Moonshot

MoonshotAI: Kimi K3

Arena Elo1558
Contender 03 · OpenAI

OpenAI: GPT-5.6 Sol

Arena Elo1560
Contender 04 · xAI

xAI: Grok 4.5

Arena Elo1564
Contender 05 · DeepSeek

DeepSeek: DeepSeek V4 Pro

Arena Elo1539

Cross-Model Capability Matrix

v2.04.14

Best value in each row highlighted in accent. Blank cells mean the provider hasn't reported that benchmark yet.

BenchmarkAnthropic: Claude Fable 5MoonshotAI: Kimi K3OpenAI: GPT-5.6 SolxAI: Grok 4.5DeepSeek: DeepSeek V4 Pro
Arena Elo
Human blind-vote Elo from LMArena — the closest proxy to 'which model feels smarter'.
15721558156015641539
AA Index
Artificial Analysis' composite Intelligence Index across reasoning, math and coding.
72.071.074.070.069.0
LiveBench
Contamination-resistant benchmark refreshed monthly — a strong signal of raw capability.
84.683.882.482.981.9
GPQA
Graduate-level science QA. Above 70% is frontier reasoning territory.
91.2%90.4%84.1%89.7%88.9%
MMLU-Pro
Broad knowledge across 57 subjects — the classic capability floor.
91.4%90.9%91.5%90.1%89.6%
SWE-Bench
Real-world GitHub bug-fixing. The single best proxy for agentic coding.
82.1%80.3%72.4%78.8%77.4%
Terminal-Bench
End-to-end shell tasks. Measures tool-use and long-horizon execution.
59.8%57.2%58.2%56.1%55.5%
ARC-AGI
Abstract reasoning puzzles. Historically brutal — even small gains matter.
23.8%22.7%46.0%22.1%20.8%
Aider Polyglot
Multi-language code editing benchmark. Signals practical dev-loop quality.
88.4%86.9%78.1%85.6%84.8%
Context
Maximum input tokens. Bigger unlocks whole-codebase and long-doc workflows.
1.0M1.0M1.1M500K1.0M
Speed (tok/s)
Output tokens per second. Matters for interactive UIs and long generations.
4255927691
Input $/1M
Cost per million input tokens.
$10.00$3.00$5.00$2.00$0.43
Output $/1M
Cost per million output tokens.
$50.00$15.00$30.00$6.00$0.87

Analysis // Model by model

Entry 01 · Anthropic

Anthropic: Claude Fable 5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Key Telemetry
Arena Elo
1572
GPQA
91.2%
SWE-Bench
82.1%
Context
1.0M
Entry 02 · Moonshot

MoonshotAI: Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Key Telemetry
Arena Elo
1558
GPQA
90.4%
SWE-Bench
80.3%
Context
1.0M
Entry 03 · OpenAI

OpenAI: GPT-5.6 Sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Key Telemetry
Arena Elo
1560
GPQA
84.1%
SWE-Bench
72.4%
Context
1.1M
Entry 04 · xAI

xAI: Grok 4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Key Telemetry
Arena Elo
1564
GPQA
89.7%
SWE-Bench
78.8%
Context
500K
Entry 05 · DeepSeek

DeepSeek: DeepSeek V4 Pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Key Telemetry
Arena Elo
1539
GPQA
88.9%
SWE-Bench
77.4%
Context
1.0M

Decision Protocol // Which to pick

01

Agentic coding

Optimize for SWE-Bench and Terminal-Bench. Aider Polyglot is the tiebreaker for real dev-loop quality.

02

Hard reasoning & research

GPQA and ARC-AGI matter more than MMLU. A high Arena Elo helps for open-ended prompts.

03

Long-context work

Filter by context window first (200K+ for full-repo work), then compare LiveBench to avoid quality drop-off.

04

Price-sensitive production

Input and output $/1M dominate at scale. Sort by price, then take the highest AA Index still in budget.