DeepSeek model benchmarks.

Compare DeepSeek models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1DeepSeek: DeepSeek V4 Pro44.3
  2. 2DeepSeek: DeepSeek V4 Flash40.3
  3. 3DeepSeek: DeepSeek V3.232

Response speed

Fastest output in this sample

  1. 1DeepSeek: DeepSeek V4 Flash92 t/s
  2. 2DeepSeek: DeepSeek V4 Pro87 t/s
  3. 3DeepSeek: DeepSeek V3.268 t/s

Input price

Lowest price per 1M tokens

  1. 1DeepSeek: DeepSeek V4 Flash$0.14/1M
  2. 2DeepSeek: DeepSeek V3.2$0.27/1M
  3. 3DeepSeek: DeepSeek V4 Pro$0.43/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare DeepSeek model capability.

Performance and price

Compare practical tradeoffs.

All DeepSeek models.

3 models in the current catalogue.

DeepSeek technical benchmarks.

Compare named evaluation records available for DeepSeek models. Missing results remain blank rather than being estimated.

BenchmarkAreaDeepSeek V4 ProDeepSeek V4 FlashDeepSeek V3.2Source
Intelligence IndexoverallQuality44.340.332OpenRouter / artificial-analysis
Coding Indexcoding59.456.244.2OpenRouter / artificial-analysis
Agentic Indexagentic36.431.118.3OpenRouter / artificial-analysis
GPQAGPQA Diamond88.889.484OpenRouter model benchmarks
Humanity's Last ExamHLE35.932.122.2OpenRouter model benchmarks
IFBenchIFBench76.579.260.7OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom96.29590.6OpenRouter model benchmarks
AA-LCRAA-LCR66.36365OpenRouter model benchmarks
GDPval-AAGDPval-AA40.334.318.5OpenRouter model benchmarks
CritPtCritPt12.97.12.9OpenRouter model benchmarks
SciCodeSciCode5044.938.9OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard46.235.635.6OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy43.337.233.5OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate64.218.3OpenRouter model benchmarks
MMLU-ProKnowledge and reasoning85OpenEvals/leaderboard-data
SWE-bench ProCoding15.56OpenEvals/leaderboard-data
SWE-bench VerifiedCoding70OpenEvals/leaderboard-data
Terminal-BenchAgentic coding39.6OpenEvals/leaderboard-data
AIME 2026Math94.17OpenEvals/leaderboard-data