System Analysis // Reference ID: 294-C

Anthropic: Claude Fable 5.1
vs the field

A forensic evaluation of today's LLM leaderboard. We pressure-test reasoning density, coding, and instruction adherence across five Tier-1 AI models — pulled live from the same aggregated feed the homepage uses.

The contenders

Primary Subject

Anthropic: Claude Fable 5.1

State of the ArtAnthropictextimagefile

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Contender 02 · Alibaba

Qwen: Qwen3.8 Max (0902)

Arena Elo1619
Contender 03 · xAI

SpaceXAI: Grok 4.6

Arena Elo1627
Contender 04 · Moonshot

MoonshotAI: Kimi K3

Arena Elo1608
Contender 05 · Z.ai

Z.ai: GLM 5.3

Arena Elo1604

Cross-Model Capability Matrix

v2.04.14

Best value in each row highlighted in accent. Blank cells mean the provider hasn't reported that benchmark yet.

BenchmarkAnthropic: Claude Fable 5.1Qwen: Qwen3.8 Max (0902)SpaceXAI: Grok 4.6MoonshotAI: Kimi K3Z.ai: GLM 5.3
Arena Elo
Human blind-vote Elo from LMArena — the closest proxy to 'which model feels smarter'.
16381619162716081604
AA Index
Artificial Analysis' composite Intelligence Index across reasoning, math and coding.
71.068.068.067.066.0
LiveBench
Contamination-resistant benchmark refreshed monthly — a strong signal of raw capability.
91.189.489.788.587.9
GPQA
Graduate-level science QA. Above 70% is frontier reasoning territory.
93.7%91.5%92.6%90.8%89.7%
MMLU-Pro
Broad knowledge across 57 subjects — the classic capability floor.
91.2%90.8%90.5%89.9%89.1%
SWE-Bench
Real-world GitHub bug-fixing. The single best proxy for agentic coding.
86.1%78.7%80.3%79.4%77.8%
Terminal-Bench
End-to-end shell tasks. Measures tool-use and long-horizon execution.
77.4%69.3%71.8%70.2%70.6%
ARC-AGI
Abstract reasoning puzzles. Historically brutal — even small gains matter.
43.8%39.7%42.6%37.2%34.8%
Aider Polyglot
Multi-language code editing benchmark. Signals practical dev-loop quality.
89.6%84.9%85.1%86.8%86.3%
Context
Maximum input tokens. Bigger unlocks whole-codebase and long-doc workflows.
1.0M1.0M500K1.0M1.3M
Speed (tok/s)
Output tokens per second. Matters for interactive UIs and long generations.
63731165871
Input $/1M
Cost per million input tokens.
$10.00$2.00$2.00$3.00$1.40
Output $/1M
Cost per million output tokens.
$50.00$6.00$6.00$15.00$4.40

Analysis // Model by model

Entry 01 · Anthropic

Anthropic: Claude Fable 5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Key Telemetry
Arena Elo
1638
GPQA
93.7%
SWE-Bench
86.1%
Context
1.0M
Entry 02 · Alibaba

Qwen: Qwen3.8 Max (0902)

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Key Telemetry
Arena Elo
1619
GPQA
91.5%
SWE-Bench
78.7%
Context
1.0M
Entry 03 · xAI

SpaceXAI: Grok 4.6

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Key Telemetry
Arena Elo
1627
GPQA
92.6%
SWE-Bench
80.3%
Context
500K
Entry 04 · Moonshot

MoonshotAI: Kimi K3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Key Telemetry
Arena Elo
1608
GPQA
90.8%
SWE-Bench
79.4%
Context
1.0M
Entry 05 · Z.ai

Z.ai: GLM 5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Key Telemetry
Arena Elo
1604
GPQA
89.7%
SWE-Bench
77.8%
Context
1.3M

Decision Protocol // Which to pick

01

Agentic coding

Optimize for SWE-Bench and Terminal-Bench. Aider Polyglot is the tiebreaker for real dev-loop quality.

02

Hard reasoning & research

GPQA and ARC-AGI matter more than MMLU. A high Arena Elo helps for open-ended prompts.

03

Long-context work

Filter by context window first (200K+ for full-repo work), then compare LiveBench to avoid quality drop-off.

04

Price-sensitive production

Input and output $/1M dominate at scale. Sort by price, then take the highest AA Index still in budget.