Quality
Highest overall scores
- 1
DeepSeek: DeepSeek V4 Pro44.3
- 2
DeepSeek: DeepSeek V4 Flash40.3
- 3
DeepSeek: DeepSeek V3.232
Compare DeepSeek models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| DeepSeek: DeepSeek V4 Pro | 44.3 | 87 t/s | 1M | $0.43 | $0.87 | |
| DeepSeek: DeepSeek V4 Flash | 40.3 | 92 t/s | 1M | $0.14 | $0.28 | |
| DeepSeek: DeepSeek V3.2 | 32 | 68 t/s | 163.8K | $0.27 | $0.40 |
Compare named evaluation records available for DeepSeek models. Missing results remain blank rather than being estimated.
| Benchmark | Area | DeepSeek V4 Pro | DeepSeek V4 Flash | DeepSeek V3.2 | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 44.3 | 40.3 | 32 | OpenRouter / artificial-analysis |
| Coding Index | coding | 59.4 | 56.2 | 44.2 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 36.4 | 31.1 | 18.3 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 88.8 | 89.4 | 84 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 35.9 | 32.1 | 22.2 | OpenRouter model benchmarks |
| IFBench | IFBench | 76.5 | 79.2 | 60.7 | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | 96.2 | 95 | 90.6 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 66.3 | 63 | 65 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 40.3 | 34.3 | 18.5 | OpenRouter model benchmarks |
| CritPt | CritPt | 12.9 | 7.1 | 2.9 | OpenRouter model benchmarks |
| SciCode | SciCode | 50 | 44.9 | 38.9 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | 46.2 | 35.6 | 35.6 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 43.3 | 37.2 | 33.5 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 6 | 4.2 | 18.3 | OpenRouter model benchmarks |
| MMLU-Pro | Knowledge and reasoning | — | — | 85 | OpenEvals/leaderboard-data |
| SWE-bench Pro | Coding | — | — | 15.56 | OpenEvals/leaderboard-data |
| SWE-bench Verified | Coding | — | — | 70 | OpenEvals/leaderboard-data |
| Terminal-Bench | Agentic coding | — | — | 39.6 | OpenEvals/leaderboard-data |
| AIME 2026 | Math | — | — | 94.17 | OpenEvals/leaderboard-data |