Moonshot AI model benchmarks.

Compare Moonshot AI models across answer quality, coding, complex tasks, speed, price, and context size.

Quality

Highest overall scores

  1. 1MoonshotAI: Kimi K357.1
  2. 2MoonshotAI: Kimi K2.7 Code41.9
  3. 3MoonshotAI: Kimi K2 Thinking17.3

Response speed

Fastest output in this sample

  1. 1MoonshotAI: Kimi K2.7 Code168 t/s
  2. 2MoonshotAI: Kimi K2 Thinking79 t/s
  3. 3MoonshotAI: Kimi K373 t/s

Input price

Lowest price per 1M tokens

  1. 1MoonshotAI: Kimi K2 Thinking$0.60/1M
  2. 2MoonshotAI: Kimi K2.7 Code$0.73/1M
  3. 3MoonshotAI: Kimi K3$3.00/1M

Highlights

Quality

Overall capability · Higher is better

Speed

Output tokens per second · Higher is better

Input price

USD per 1M tokens · Lower is better

Quality benchmarks

Compare Moonshot AI model capability.

Performance and price

Compare practical tradeoffs.

All Moonshot AI models.

3 models in the current catalogue.

ModelLinks
MoonshotAI: Kimi K357.173 t/s1M$3.00$15.00
MoonshotAI: Kimi K2.7 Code41.9168 t/s262.1K$0.73$3.50
MoonshotAI: Kimi K2 Thinking17.379 t/s262.1K$0.60$2.50

Moonshot AI technical benchmarks.

Compare named evaluation records available for Moonshot AI models. Missing results remain blank rather than being estimated.

BenchmarkAreaKimi K3Kimi K2.7 CodeKimi K2 ThinkingSource
Intelligence IndexoverallQuality57.141.917.3OpenRouter / artificial-analysis
Coding Indexcoding76.260.821OpenRouter / artificial-analysis
Agentic Indexagentic50.129.61.8OpenRouter / artificial-analysis
GPQAGPQA Diamond93.589.671.3OpenRouter model benchmarks
Humanity's Last ExamHLE44.332.89.5OpenRouter model benchmarks
AA-LCRAA-LCR74.766.352.7OpenRouter model benchmarks
GDPval-AAGDPval-AA59.434.40OpenRouter model benchmarks
CritPtCritPt23.4100OpenRouter model benchmarks
SciCodeSciCode58.747.533OpenRouter model benchmarks
AA-Omniscience AccuracyAA-Omniscience Accuracy4638.615.7OpenRouter model benchmarks
AA-Omniscience Non-Hallucination RateAA-Omniscience Non-Hallucination Rate49.119.741.1OpenRouter model benchmarks
IFBenchIFBench63.162.8OpenRouter model benchmarks
τ²-Bench Telecomτ²-Bench Telecom90.125.4OpenRouter model benchmarks
Terminal-Bench HardTerminal-Bench Hard44.76.8OpenRouter model benchmarks
MMLU-ProKnowledge and reasoning84.6OpenEvals/leaderboard-data
SWE-bench VerifiedCoding71.3OpenEvals/leaderboard-data
Terminal-BenchAgentic coding35.7OpenEvals/leaderboard-data