OpenAI model benchmarks.

Compare OpenAI models across answer quality, coding, complex tasks, speed, price, and context size.

OpenAI

Quality

Highest overall scores

  1. 1OpenAI: GPT-5.6 Sol58.9
  2. 2OpenAI GPT Mini Latest0
  3. 3OpenAI: GPT-5.6 Luna Pro0

Response speed

Fastest output in this sample

  1. 1OpenAI: GPT-5.6 Luna Pro170 t/s
  2. 2OpenAI: GPT-5.6 Sol85 t/s
  3. 3OpenAI GPT Mini Latest0 t/s

Input price

Lowest price per 1M tokens

  1. 1OpenAI: GPT-5.6 Luna Pro$0.50/1M
  2. 2OpenAI GPT Mini Latest$0.75/1M
  3. 3OpenAI: GPT-5.6 Sol$5.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare OpenAI model capability.

Performance and price

Compare practical tradeoffs.

All OpenAI models.

3 models in the current catalogue.

OpenAI technical benchmarks.

Compare named evaluation records available for OpenAI models. Missing results remain blank rather than being estimated.

BenchmarkAreaGPT-5.6 SolOpenAI GPT Mini LatestGPT-5.6 Luna ProSource
Intelligence IndexoverallQuality58.9OpenRouter / artificial-analysis
Coding Indexcoding77.4OpenRouter / artificial-analysis
Agentic Indexagentic54OpenRouter / artificial-analysis
GPQAGPQA Diamond94.1OpenRouter model benchmarks
Humanity's Last ExamHLE47.2OpenRouter model benchmarks
IFBenchIFBench72.7OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom85.1OpenRouter model benchmarks
AA-LCRAA-LCR73.7OpenRouter model benchmarks
GDPval-AAGDPval-AA61.8OpenRouter model benchmarks
CritPtCritPt32.3OpenRouter model benchmarks
SciCodeSciCode56.1OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard65.9OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy58.5OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate11.2OpenRouter model benchmarks