Model evidence profile

Kimi K2.5

Moonshot AI · Open Weight · Non-Reasoning

currentsupported
Overall public score
59.14
Source rank
#82
Evidence coverage
9 benchmarks · 2 sources

Strongest published evidencePublic overall score 59.14 at source rank #82.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingInstruction FollowingKnowledgeMathMultilingualMultimodal groundedOverallReasoning
Agentic percentile
18.0 percentileRank #115 of 140
Coding percentile
69.4 percentileRank #45 of 145
Instruction Following percentile
80.5 percentileRank #9 of 42
Knowledge percentile
14.8 percentileRank #47 of 55
Math percentile
33.3 percentileRank #5 of 7
Multilingual percentile
41.7 percentileRank #8 of 13
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticsupported
39.5
Rank
#115 of 140
Percentile
18.0%
Benchmarks
9
Codingsupported
56.8
Rank
#45 of 145
Percentile
69.4%
Benchmarks
9
Instruction Followingsupported
91.2
Rank
#9 of 42
Percentile
80.5%
Benchmarks
9
Knowledgesupported
54.6
Rank
#47 of 55
Percentile
14.8%
Benchmarks
9
Mathsupported
62.5
Rank
#5 of 7
Percentile
33.3%
Benchmarks
9
Multilingualsupported
38.2
Rank
#8 of 13
Percentile
41.7%
Benchmarks
9
Overallsupported
59.1
Rank
#82
Percentile
Unavailable
Benchmarks
9
Reasoningsupported
44.9
Rank
Not ranked
Percentile
Unavailable
Benchmarks
9

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

moonshot-aibenchlm:kimi-k2-5
primary
Input / 1M
$0.60
Cached input / 1M
Unavailable
Output / 1M
$3.00
Context
256,000
View price source
Context window
256,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-02-01
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
39.47#115Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
56.80#45Not publishedAug 29, 2026BenchLM

instructionFollowing

ScoreRankWeightLast UpdatedSource
91.20#9Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
54.60#47Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
62.50#5Not publishedAug 29, 2026BenchLM

multilingual

ScoreRankWeightLast UpdatedSource
38.20#8Not publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
60.70Not rankedNot publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
44.90Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
59.14#82Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model