AI model benchmarks.

Compare broadly useful AI models across intelligence, coding, tool use, response speed, price, context size, and availability. Every section explains what is measured and whether a higher or lower value is better.

Intelligence

Strongest performance across published evaluations

  1. 1Claude Opus 560.7
  2. 2OpenAI: GPT-5.6 Sol58.9

Response speed

Fastest output in this sample

  1. 1MiniMax: MiniMax M2.7294.5 t/s
  2. 2NVIDIA: Nemotron 3 Nano 30B A3B189 t/s

Cost per task

Lowest estimated cost for a typical task

  1. 1IBM: Granite 4.0 Micro$0.0001
  2. 2Qwen: Qwen3.7 Flash$0.0001

Context size

Most content read at once

  1. 1OpenAI: GPT-5.6 Sol1.1M
  2. 2Xiaomi: MiMo-V2.5-Pro1.1M

Highlights

Intelligence

Published evaluation score · Higher is better

Output speed

Generated tokens per second · Higher is better

Cost per task

Estimated cost for a typical task · Lower is better

Intelligence

How capable is each model?

Intelligence combines published evaluations into a comparable score. Use the tabs to focus on general capability, writing code, or completing tasks with tools. Higher is better.

Named evaluations

Explore technical benchmarks.

Select an evaluation such as GPQA, MMLU-Pro, Humanity's Last Exam, SWE-bench, Terminal-Bench, or AIME. Only models with a compatible published result are ranked.

RankModelCompanyScoreSource
1OpenAI: GPT-5.6 SolOpenAI94.1OpenRouter model benchmarks
2MoonshotAI: Kimi K3Moonshot AI93.5OpenRouter model benchmarks
3Claude Opus 5Anthropic93.2OpenRouter model benchmarks
4xAI: Grok 4.5xAI93.1OpenRouter model benchmarks
5MiniMax: MiniMax M3MiniMax92.9OpenRouter model benchmarks
6Google: Gemini 3.6 FlashGoogle92.8OpenRouter model benchmarks
7Qwen: Qwen3.7 MaxAlibaba92.3OpenRouter model benchmarks
8Google: Gemini 3.5 FlashGoogle92.2OpenRouter model benchmarks
9Anthropic: Claude Sonnet 5Anthropic91.1OpenRouter model benchmarks
10Qwen: Qwen3.7 PlusAlibaba90OpenRouter model benchmarks
11Meta: Muse Spark 1.1Meta89.8OpenRouter model benchmarks
12Tencent: Hy3 previewTencent89.7OpenRouter model benchmarks
13MoonshotAI: Kimi K2.7 CodeMoonshot AI89.6OpenRouter model benchmarks
14Z.ai: GLM 5.2Z.ai89.5OpenRouter model benchmarks
15xAI: Grok Build 0.1xAI89.5OpenRouter model benchmarks
16DeepSeek: DeepSeek V4 FlashDeepSeek89.4OpenRouter model benchmarks
17DeepSeek: DeepSeek V4 ProDeepSeek88.8OpenRouter model benchmarks
18MiniMax: MiniMax M2.7MiniMax87.4OpenRouter model benchmarks
19Thinking Machines: InklingThinking Machines87.2OpenRouter model benchmarks
20Z.ai: GLM 5.1Z.ai86.8OpenRouter model benchmarks
21Xiaomi: MiMo-V2.5-ProXiaomi86.6OpenRouter model benchmarks
22MiniMax: MiniMax M2.5MiniMax85.2OpenEvals/leaderboard-data
23Xiaomi: MiMo-V2.5Xiaomi84.9OpenRouter model benchmarks
24DeepSeek: DeepSeek V3.2DeepSeek84OpenRouter model benchmarks
25StepFun: Step 3.5 FlashStepFun83.1OpenRouter model benchmarks
26Amazon: Nova 2 LiteAmazon81.1OpenRouter model benchmarks
27StepFun: Step 3.7 FlashStepFun80.9OpenRouter model benchmarks
28NVIDIA: Nemotron 3 SuperNVIDIA80OpenRouter model benchmarks
29Mistral: Mistral Small 4Mistral AI76.9OpenRouter model benchmarks
30Cohere: Command ACohere76.1OpenRouter model benchmarks
31NVIDIA: Nemotron 3 Nano 30B A3BNVIDIA75.7OpenRouter model benchmarks
32Z.ai: GLM 4.7 FlashZ.ai75.2OpenEvals/leaderboard-data
33Mistral: Mistral Medium 3.5Mistral AI74.8OpenRouter model benchmarks
34Upstage: Solar Pro 3Upstage72.4OpenRouter model benchmarks
35MoonshotAI: Kimi K2 ThinkingMoonshot AI71.3OpenRouter model benchmarks
36Mistral: Mistral Large 3 2512Mistral AI68OpenRouter model benchmarks
37Meta: Llama 4 MaverickMeta67.1OpenRouter model benchmarks
38Perplexity: Sonar ProPerplexity57.8OpenRouter model benchmarks
39Microsoft: Phi 4Microsoft57.5OpenRouter model benchmarks
40Amazon: Nova Premier 1.0Amazon56.9OpenRouter model benchmarks
41Reka Flash 3Reka AI52.9OpenRouter model benchmarks
42Meta: Llama 3.3 70B InstructMeta49.8OpenRouter model benchmarks
43Amazon: Nova Lite 1.0Amazon43.3OpenRouter model benchmarks
44IBM: Granite 4.1 8BIBM43.3OpenRouter model benchmarks
45AI21: Jamba Large 1.7AI21 Labs39OpenRouter model benchmarks
46IBM: Granite 4.0 MicroIBM30.3OpenRouter model benchmarks

Intelligence comparisons

Compare capability with practical tradeoffs.

Start with intelligence, then inspect the practical measurement that matters to you. Lower is better for cost and answer-start time; higher is better for speed and context.

Token use

How long can one answer be?

This shows the maximum response length allowed by each model, not the number of tokens a typical answer will actually use. Higher means the model can produce a longer single response.

Cost

What does using each model cost?

Cost per task estimates 1,000 input tokens and 500 output tokens. Input and output prices show the underlying API price per one million tokens. Lower is better.

Context window

How much information can the model read at once?

The context window is the combined amount of instructions, conversation, documents, and generated text the model can consider in one request. Higher is better for large documents and long conversations.

Output speed

How quickly does the answer appear?

Output speed measures how many tokens arrive each second after the model begins responding. Higher is better. These are observed OpenRouter endpoint results, not a guarantee for every provider or request.

Answer start time

How long before the model starts replying?

This is the delay before the first part of an answer arrives. Lower is better. A model can start quickly but still generate the rest of its answer slowly.

Total response time

How long for a complete typical answer?

This estimate combines answer-start time with the time required to generate 500 output tokens. Lower is better. Actual time changes with answer length, provider load, and reasoning settings.

Model size

How large are open-weight models?

Parameter count is one rough indicator of the memory and hardware needed to run a model locally. It is not an intelligence score, and providers do not publish it for every model.

Prompt options

What can developers control?

These are common API controls supported by models in this catalogue. Availability can vary by provider. The count shows how many listed models support each option.

Max Tokens

Supported by 67 of 67 listed models

Temperature

Supported by 62 of 67 listed models

Top P

Supported by 58 of 67 listed models

Tools

Supported by 56 of 67 listed models

Stop

Supported by 54 of 67 listed models

Tool Choice

Supported by 54 of 67 listed models

Include Reasoning

Supported by 52 of 67 listed models

Reasoning

Supported by 52 of 67 listed models

Response Format

Supported by 52 of 67 listed models

Structured Outputs

Supported by 52 of 67 listed models

Frequency Penalty

Supported by 47 of 67 listed models

Seed

Supported by 47 of 67 listed models