Head-to-head · live
GPT vs Claude
Side-by-side comparison of the current flagship models from OpenAI and Anthropic across every benchmark that matters — refreshed live, no marketing spin.
OpenAI
OpenAI: GPT-5.6 SolAnthropic
Anthropic: Claude Fable 5Row-by-row
| Benchmark | OpenAI: GPT-5.6 Sol | Anthropic: Claude Fable 5 | Winner |
|---|---|---|---|
| Arena Elo | 1560 | 1572 | Claude |
| AA Index | 74.0 | 72.0 | GPT |
| LiveBench | 82.4 | 84.6 | Claude |
| GPQA | 84.1% | 91.2% | Claude |
| SWE-bench | 72.4% | 82.1% | Claude |
| Terminal-Bench | 58.2% | 59.8% | Claude |
| Aider Polyglot | 78.1% | 88.4% | Claude |
| Output $/1M | $30.00 | $50.00 | GPT |
Frequently asked
- GPT vs Claude — which is better?
- It depends on the task. Claude flagships (Opus/Fable/Mythos) consistently lead on SWE-bench and long-form writing quality; GPT flagships (Sol/Terra) lead on raw reasoning benchmarks and multimodality. The tables above show the current numbers head-to-head.
- Is Claude smarter than GPT?
- On some benchmarks yes, on others no. Recent Claude flagships lead on SWE-bench Verified (coding) and Terminal-Bench (agentic), while OpenAI's Sol line leads on GPQA and ARC-AGI. See the row-by-row comparison above.
- Which is cheaper, GPT or Claude?
- GPT-family flagships and Claude flagships are priced within ~30% of each other for output tokens. Smaller siblings (mini, haiku, flash) are 10× cheaper on both sides. See /cheapest for the full price ranking.
Also see: best LLM, best coding LLM, top-5 comparison.