#76DeepSeek

DeepSeek 4.1 Flash.

Context
—
Input $/1M
—
Output $/1M
—
Modality
text
Arena Elo
—
LMArena — human blind votes
AA Index
—
Artificial Analysis composite
LiveBench
—
Contamination-resistant
GPQA
—
Grad-level science reasoning
MMLU-Pro
—
Broad knowledge
SWE-bench
—
Real GitHub bugs (Verified)
Terminal-Bench
—
Agentic shell tasks
ARC-AGI
—
Abstract reasoning
Aider Polyglot
—
Multi-language code editing
Speed
—
Median output speed

Profile

DeepSeek 4.1 Flash is DeepSeek's entry at position #76 on the current aggregated leaderboard. Its benchmark profile is still being filled in as independent evaluations publish results. Every figure below is aggregated from the public sources listed on the homepage and refreshed on each pipeline run.

Caveats worth knowing before you read too much into the position: no execution-graded coding result is public yet; it has not accumulated enough Arena votes for a stable Elo; independent throughput measurements are not yet available. As with every model on this index, treat small gaps as noise and test on your own workload before committing.

How this model's position is computed is documented in the methodology; what each benchmark actually measures is in the benchmark guide. See a figure that disagrees with its source? Report it — accepted corrections apply on the next refresh.

Nearby on the leaderboard