About
Why this index exists
SOTA Model answers one question that is surprisingly hard to answer honestly: which large language model is actually the best right now, and how sure can you be about that?
Last reviewed · August 2026
The problem we set out to solve
The frontier moves weekly. A model that led every chart in March is mid-pack by June, and the places you would go to check — a lab's own launch post, a crowded Arena leaderboard, a benchmark table on someone's blog — each answer a slightly different question, on a different date, with different incentives. Launch posts are marketing. Arena rewards answers that read well. Static benchmarks leak into training sets and quietly stop measuring anything. Private evaluations cannot be reproduced by the person reading them.
Nobody is lying, exactly. But if you are choosing a model for production work, or you simply want to know what "state of the art" means this week, you end up opening nine tabs and doing the reconciliation yourself. This site is that reconciliation, done continuously and shown with its sources attached.
What we actually publish
- A single current SOTA pick, with the reasoning and benchmark spread that produced it — not a vague "top models" list that avoids committing.
- A live ranked table of frontier models with per-benchmark scores: Arena Elo, GPQA Diamond, SWE-bench Verified, Terminal-Bench, ARC-AGI-2, Aider Polyglot, LiveBench, MMLU-Pro, price and throughput. Every column is sortable, because your priorities are not our priorities.
- Use-case leaderboards — coding, reasoning, agentic, context length, speed, cost — because "best model" is meaningless without a job attached to it.
- Release and rumor tracking, so you can see how new the current leader is and what is credibly reported to be landing next.
- A news timeline filtered to model releases, capability results and lab announcements, rather than general AI commentary.
Editorial principles
- Show the source. Every score links out to where it came from. If we cannot point at a public number, we mark it as an estimate rather than dressing it up.
- No single benchmark decides. A model is SOTA when it leads across the widest set of independent signals, not because it won one eval its lab happened to publish.
- Flagships compete with flagships. Mini, lite, nano, flash and air variants are tracked but are never allowed to displace true frontier models at the top, no matter how well they score on a cheap-to-game metric.
- Freshness beats polish. The pipeline runs roughly every two hours. A slightly rough number from this morning is worth more to you than a beautifully formatted one from last month.
- Say when we are unsure. Below the top handful of models, scores are estimates. We label them as such on the methodology page instead of implying false precision.
Independence and funding
SOTA Model is an independent project. It is not operated by, sponsored by, or affiliated with OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, Moonshot, Mistral, Alibaba, LMArena, Artificial Analysis, OpenRouter, or any other lab or benchmark organisation. Model names and trademarks belong to their respective owners and are used descriptively.
The site is funded by display advertising. Advertising has no influence on rankings: the ranking pipeline runs before any page is rendered and has no knowledge of advertisers. Ads on this site are served as non-personalized, so they do not rely on behavioural profiles built from your browsing. We do not take payment for placement, inclusion, or a better position in any table, and we will say so publicly if that ever changes.
What this site is not
It is not an authoritative benchmark authority. We do not run evaluations ourselves — we aggregate, normalise and reconcile public ones. It is not investment or procurement advice. And it is not a replacement for testing a model on your own workload, which remains the only evaluation that actually measures the thing you care about. Use the leaderboard to shortlist two or three candidates, then test those against your real prompts.
Corrections
If a score is wrong, a model is missing, or a ranking looks indefensible, we want to hear it — those reports are the main way this index improves. Include the model, the benchmark, and a link to the number you believe is correct, and send it via the contact page. Corrections to published scores are applied on the next refresh cycle.
To understand how a ranking was produced before you dispute it, read the methodology, and see the benchmark guide for what each evaluation does and does not measure.