xAI model benchmarks.

Compare xAI models across answer quality, coding, complex tasks, speed, price, and context size.

xAI

Quality

Highest overall scores

  1. 1xAI: Grok 4.553.8
  2. 2xAI: Grok Build 0.139.8
  3. 3xAI: Grok Latest0

Response speed

Fastest output in this sample

  1. 1xAI: Grok Build 0.1115 t/s
  2. 2xAI: Grok 4.554 t/s
  3. 3xAI: Grok Latest0 t/s

Input price

Lowest price per 1M tokens

  1. 1xAI: Grok Build 0.1$1.00/1M
  2. 2xAI: Grok 4.5$2.00/1M
  3. 3xAI: Grok Latest$2.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare xAI model capability.

Performance and price

Compare practical tradeoffs.

All xAI models.

3 models in the current catalogue.

ModelLinks
xAI: Grok 4.553.854 t/s500K$2.00$6.00
xAI: Grok Build 0.139.8115 t/s256K$1.00$2.00
xAI: Grok Latest0500K$2.00$6.00

xAI technical benchmarks.

Compare named evaluation records available for xAI models. Missing results remain blank rather than being estimated.

BenchmarkAreaGrok 4.5Grok Build 0.1Grok LatestSource
Intelligence IndexoverallQuality53.839.8OpenRouter / artificial-analysis
Coding Indexcoding72.451.5OpenRouter / artificial-analysis
Agentic Indexagentic45.728OpenRouter / artificial-analysis
GPQAGPQA Diamond93.189.5OpenRouter model benchmarks
Humanity's Last ExamHLE40.336OpenRouter model benchmarks
AA-LCRAA-LCR67.764.7OpenRouter model benchmarks
GDPval-AAGDPval-AA51.435.7OpenRouter model benchmarks
CritPtCritPt15.49.1OpenRouter model benchmarks
SciCodeSciCode54.150.2OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy5251.4OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate46.59.5OpenRouter model benchmarks