Compare AI models.
Choose between two and five models to compare their quality, speed, context size, and pricing side by side.
Choose models.
Select up to five models.
2 of 5 selected
Side-by-side comparison.
Comparing 2 models.
| Model | Company | Links | ||||||
|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | 60.7 | 63 t/s | 1M | $0.018 | $5.00 | $25.00 | ||
| GPT-5.6 Sol | 58.9 | 85 t/s | 1.1M | $0.020 | $5.00 | $30.00 |
Cost per task estimates 1,000 input tokens and 500 output tokens using the listed API prices.
Named benchmark comparison.
Compare GPQA, MMLU-Pro, Humanity's Last Exam, SWE-bench, Terminal-Bench, AIME, and other available evaluation records directly.
| Benchmark | Area | Claude Opus 5 | GPT-5.6 Sol | Source |
|---|---|---|---|---|
| Intelligence Index | overallQuality | 60.7 | 58.9 | OpenRouter / artificial-analysis |
| Coding Index | coding | 78 | 77.4 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 55.3 | 54 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 93.2 | 94.1 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 52.6 | 47.2 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 70 | 73.7 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 68.1 | 61.8 | OpenRouter model benchmarks |
| CritPt | CritPt | 29.1 | 32.3 | OpenRouter model benchmarks |
| SciCode | SciCode | 55.7 | 56.1 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 54.2 | 58.5 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 49.9 | 11.2 | OpenRouter model benchmarks |
| IFBench | IFBench | — | 72.7 | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | — | 85.1 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | — | 65.9 | OpenRouter model benchmarks |