Model evidence profile

GPT-5.1

OpenAI · Proprietary · Reasoning

currentestimated
Overall public score
53.66
Source rank
#112
Evidence coverage
3 benchmarks · 2 sources

Strongest published evidencePublic overall score 53.66 at source rank #112.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
49.6 percentileRank #71 of 140
Coding percentile
59.7 percentileRank #59 of 145
Knowledge percentile
Knowledge: Unavailable
Math percentile
Math: Unavailable
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
49.0
Rank
#71 of 140
Percentile
49.6%
Benchmarks
3
Codingestimated
53.4
Rank
#59 of 145
Percentile
59.7%
Benchmarks
3
Mathestimated
49.7
Rank
Not ranked
Percentile
Unavailable
Benchmarks
3
Overallestimated
53.7
Rank
#112
Percentile
Unavailable
Benchmarks
3

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

openaibenchlm:gpt-5-1
primary
Input / 1M
$1.25
Cached input / 1M
Unavailable
Output / 1M
$10.00
Context
400,000
View price source
Context window
200,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2025-11-13
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
48.98#71Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
53.44#59Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
49.70Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
53.66#112Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model