Quality
Highest overall scores
- 1
Mistral: Mistral Medium 3.529.9
- 2
Mistral: Mistral Small 419.6
- 3
Mistral: Mistral Large 3 251215.9
Compare Mistral AI models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Mistral: Mistral Medium 3.5 | 29.9 | 35 t/s | 262.1K | $1.50 | $7.50 | |
| Mistral: Mistral Small 4 | 19.6 | 89 t/s | 262.1K | $0.15 | $0.60 | |
| Mistral: Mistral Large 3 2512 | 15.9 | 44 t/s | 262.1K | $0.50 | $1.50 |
Compare named evaluation records available for Mistral AI models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Mistral Medium 3.5 | Mistral Small 4 | Mistral Large 3 2512 | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 29.9 | 19.6 | 15.9 | OpenRouter / artificial-analysis |
| Coding Index | coding | 46.9 | 26.6 | 20.1 | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 19 | 4.7 | 5.5 | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 74.8 | 76.9 | 68 | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 12.8 | 9.5 | 4.1 | OpenRouter model benchmarks |
| IFBench | IFBench | 68.8 | 48.2 | 36.2 | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | 94.2 | 41.2 | 24.6 | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 61 | 44.7 | 34.7 | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 21.6 | 4.6 | 7 | OpenRouter model benchmarks |
| CritPt | CritPt | 0 | 0.3 | 0 | OpenRouter model benchmarks |
| SciCode | SciCode | 39.6 | 38 | 36.2 | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | 33.3 | 17.4 | 15.9 | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 25.1 | 22.1 | 24.1 | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 18 | 33.2 | 16.3 | OpenRouter model benchmarks |