Quality
Highest overall scores
- 1
Qwen: Qwen3.7 Max46
- 2
Qwen: Qwen3.7 Plus39
- 3
Qwen: Qwen3.7 Flash0
Compare Alibaba models across answer quality, coding, complex tasks, speed, price, and context size.
Highest overall scores
Fastest output in this sample
Lowest price per 1M tokens
Overall capability · Higher is better
Output tokens per second · Higher is better
USD per 1M tokens · Lower is better
Quality benchmarks
Performance and price
3 models in the current catalogue.
| Model | Links | |||||
|---|---|---|---|---|---|---|
| Qwen: Qwen3.7 Max | 46 | 35 t/s | 1M | $1.25 | $3.75 | |
| Qwen: Qwen3.7 Plus | 39 | 13 t/s | 1M | $0.32 | $1.28 | |
| Qwen: Qwen3.7 Flash | 0 | 104 t/s | 1M | $0.03 | $0.13 |
Compare named evaluation records available for Alibaba models. Missing results remain blank rather than being estimated.
| Benchmark | Area | Qwen3.7 Max | Qwen3.7 Plus | Qwen3.7 Flash | Source |
|---|---|---|---|---|---|
| Intelligence Index | overallQuality | 46 | 39 | — | OpenRouter / artificial-analysis |
| Coding Index | coding | 66 | 55.9 | — | OpenRouter / artificial-analysis |
| Agentic Index | agentic | 30.6 | 20.8 | — | OpenRouter / artificial-analysis |
| GPQA | GPQA Diamond | 92.3 | 90 | — | OpenRouter model benchmarks |
| Humanity's Last Exam | HLE | 38.1 | 33.4 | — | OpenRouter model benchmarks |
| IFBench | IFBench | 80.5 | 78 | — | OpenRouter model benchmarks |
| τ²-Bench Telecom | τ²-Bench Telecom | 94.7 | 93 | — | OpenRouter model benchmarks |
| AA-LCR | AA-LCR | 69 | 65 | — | OpenRouter model benchmarks |
| GDPval-AA | GDPval-AA | 38.6 | 22.1 | — | OpenRouter model benchmarks |
| CritPt | CritPt | 13.4 | 9.1 | — | OpenRouter model benchmarks |
| SciCode | SciCode | 48.8 | 45.5 | — | OpenRouter model benchmarks |
| Terminal-Bench Hard | Terminal-Bench Hard | 50.8 | 47 | — | OpenRouter model benchmarks |
| AA-Omniscience Accuracy | AA-Omniscience Accuracy | 30.1 | 22.2 | — | OpenRouter model benchmarks |
| AA-Omniscience Non-Hallucination Rate | AA-Omniscience Non-Hallucination Rate | 77.1 | 74.5 | — | OpenRouter model benchmarks |