Head-to-head · live

GPT vs Claude

Side-by-side comparison of the current flagship models from OpenAI and Anthropic across every benchmark that matters — refreshed live, no marketing spin.

Row-by-row

BenchmarkOpenAI: GPT-5.5 ProAnthropic: Claude Fable 5.1Winner
Arena Elo16011635Claude
AA Index70.075.0Claude
LiveBench87.990.6Claude
GPQA92.0%93.6%Claude
SWE-bench78.2%86.1%Claude
Terminal-Bench65.1%73.2%Claude
Aider Polyglot81.5%89.2%Claude
Output $/1M$180.00$50.00Claude

Frequently asked

GPT vs Claude — which is better?
It depends on the task. Claude flagships (Opus/Fable/Mythos) consistently lead on SWE-bench and long-form writing quality; GPT flagships (Sol/Terra) lead on raw reasoning benchmarks and multimodality. The tables above show the current numbers head-to-head.
Is Claude smarter than GPT?
On some benchmarks yes, on others no. Recent Claude flagships lead on SWE-bench Verified (coding) and Terminal-Bench (agentic), while OpenAI's Sol line leads on GPQA and ARC-AGI. See the row-by-row comparison above.
Which is cheaper, GPT or Claude?
GPT-family flagships and Claude flagships are priced within ~30% of each other for output tokens. Smaller siblings (mini, haiku, flash) are 10× cheaper on both sides. See /cheapest for the full price ranking.

Also see: best LLM, best coding LLM, top-5 comparison.