Quality
Highest overall scores
- 1
Claude Opus 560.7
- 2
Anthropic: Claude Sonnet 553.4
- 3
Claude Opus 5 (Fast)0
Compare Anthropic models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Claude Opus 5 | 60.7 | 63 t/s | 1M | $5.00 | $25.00 | |
| Anthropic: Claude Sonnet 5 | 53.4 | 96.5 t/s | 1M | $2.00 | $10.00 | |
| Claude Opus 5 (Fast) | 0 | 112 t/s | 1M | $10.00 | $50.00 |
Compare named evaluation records available for Anthropic models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Claude Opus 5 | Claude Sonnet 5 | Claude Opus 5 (Fast) | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 60.7 | 53.4 | — | OpenRouter / artificial-analysis |
| Coding Index | coding | 78 | 71.5 | — | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 55.3 | 46.7 | — | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 93.2 | 91.1 | — | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 52.6 | 39.6 | — | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 70 | 70.7 | — | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 68.1 | 55.2 | — | OpenRouter model benchmarks |
| CritPt | CritPt | 29.1 | 16.9 | — | OpenRouter model benchmarks |
| SciCode | SciCode | 55.7 | 53.6 | — | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 54.2 | 38.3 | — | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 49.9 | 62.7 | — | OpenRouter model benchmarks |