Head-to-head · live

GPT vs Claude

Side-by-side comparison of the current flagship models from OpenAI and Anthropic across every benchmark that matters — refreshed live, no marketing spin.

Row-by-row

BenchmarkOpenAI: GPT-5.6 SolAnthropic: Claude Fable 5Winner
Arena Elo15601572Claude
AA Index74.072.0GPT
LiveBench82.484.6Claude
GPQA84.1%91.2%Claude
SWE-bench72.4%82.1%Claude
Terminal-Bench58.2%59.8%Claude
Aider Polyglot78.1%88.4%Claude
Output $/1M$30.00$50.00GPT

Frequently asked

GPT vs Claude — which is better?
It depends on the task. Claude flagships (Opus/Fable/Mythos) consistently lead on SWE-bench and long-form writing quality; GPT flagships (Sol/Terra) lead on raw reasoning benchmarks and multimodality. The tables above show the current numbers head-to-head.
Is Claude smarter than GPT?
On some benchmarks yes, on others no. Recent Claude flagships lead on SWE-bench Verified (coding) and Terminal-Bench (agentic), while OpenAI's Sol line leads on GPQA and ARC-AGI. See the row-by-row comparison above.
Which is cheaper, GPT or Claude?
GPT-family flagships and Claude flagships are priced within ~30% of each other for output tokens. Smaller siblings (mini, haiku, flash) are 10× cheaper on both sides. See /cheapest for the full price ranking.

Also see: best LLM, best coding LLM, top-5 comparison.