Model evidence profile

Gemini 3.1 Pro

Google · Proprietary · Reasoning

currentestimated
Overall public score
56.25
Source rank
#99
Evidence coverage
7 benchmarks · 2 sources

Strongest published evidencePublic overall score 56.25 at source rank #99.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
9.4 percentileRank #127 of 140
Coding percentile
41.7 percentileRank #85 of 145
Knowledge percentile
44.4 percentileRank #31 of 55
Math percentile
Math: Unavailable
Multimodal grounded percentile
79.4 percentileRank #8 of 35
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
33.8
Rank
#127 of 140
Percentile
9.4%
Benchmarks
7
Codingestimated
49.3
Rank
#85 of 145
Percentile
41.7%
Benchmarks
7
Knowledgeestimated
66.7
Rank
#31 of 55
Percentile
44.4%
Benchmarks
7
Mathestimated
55.1
Rank
Not ranked
Percentile
Unavailable
Benchmarks
7
Overallestimated
56.3
Rank
#99
Percentile
Unavailable
Benchmarks
7
Reasoningestimated
72.4
Rank
Not ranked
Percentile
Unavailable
Benchmarks
7

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

googlebenchlm:gemini-3-1-pro
primary
Input / 1M
$2.00
Cached input / 1M
$0.20
Output / 1M
$12.00
Context
1,000,000
View price source
Context window
1,000,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-02-19
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
33.84#127Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
49.30#85Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
66.70#31Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
55.10Not rankedNot publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
76.80#8Not publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
72.40Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
56.25#99Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model