Intelligence
Strongest performance across published evaluations
- 1
Claude Opus 560.7
- 2
OpenAI: GPT-5.6 Sol58.9
Compare broadly useful AI models across intelligence, coding, tool use, response speed, price, context size, and availability. Every section explains what is measured and whether a higher or lower value is better.
Strongest performance across published evaluations
Fastest output in this sample
Lowest estimated cost for a typical task
Most content read at once
Published evaluation score · Higher is better
Generated tokens per second · Higher is better
Estimated cost for a typical task · Lower is better
Intelligence
Intelligence combines published evaluations into a comparable score. Use the tabs to focus on general capability, writing code, or completing tasks with tools. Higher is better.
Named evaluations
Select an evaluation such as GPQA, MMLU-Pro, Humanity's Last Exam, SWE-bench, Terminal-Bench, or AIME. Only models with a compatible published result are ranked.
Intelligence comparisons
Start with intelligence, then inspect the practical measurement that matters to you. Lower is better for cost and answer-start time; higher is better for speed and context.
Token use
This shows the maximum response length allowed by each model, not the number of tokens a typical answer will actually use. Higher means the model can produce a longer single response.
Cost
Cost per task estimates 1,000 input tokens and 500 output tokens. Input and output prices show the underlying API price per one million tokens. Lower is better.
Context window
The context window is the combined amount of instructions, conversation, documents, and generated text the model can consider in one request. Higher is better for large documents and long conversations.
Output speed
Output speed measures how many tokens arrive each second after the model begins responding. Higher is better. These are observed OpenRouter endpoint results, not a guarantee for every provider or request.
Answer start time
This is the delay before the first part of an answer arrives. Lower is better. A model can start quickly but still generate the rest of its answer slowly.
Total response time
This estimate combines answer-start time with the time required to generate 500 output tokens. Lower is better. Actual time changes with answer length, provider load, and reasoning settings.
Model size
Parameter count is one rough indicator of the memory and hardware needed to run a model locally. It is not an intelligence score, and providers do not publish it for every model.
Prompt options
These are common API controls supported by models in this catalogue. Availability can vary by provider. The count shows how many listed models support each option.
Supported by 67 of 67 listed models
Supported by 62 of 67 listed models
Supported by 58 of 67 listed models
Supported by 56 of 67 listed models
Supported by 54 of 67 listed models
Supported by 54 of 67 listed models
Supported by 52 of 67 listed models
Supported by 52 of 67 listed models
Supported by 52 of 67 listed models
Supported by 52 of 67 listed models
Supported by 47 of 67 listed models
Supported by 47 of 67 listed models