Model evidence profile

GPT-5.4 mini

OpenAI · Proprietary · Reasoning

currentestimated
Overall public score
56.96
Source rank
#94
Evidence coverage
6 benchmarks · 2 sources

Strongest published evidencePublic overall score 56.96 at source rank #94.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
14.4 percentileRank #120 of 140
Coding percentile
29.9 percentileRank #102 of 145
Knowledge percentile
18.5 percentileRank #45 of 55
Math percentile
Math: Unavailable
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
37.5
Rank
#120 of 140
Percentile
14.4%
Benchmarks
6
Codingestimated
46.6
Rank
#102 of 145
Percentile
29.9%
Benchmarks
6
Knowledgeestimated
57.3
Rank
#45 of 55
Percentile
18.5%
Benchmarks
6
Mathestimated
44.6
Rank
Not ranked
Percentile
Unavailable
Benchmarks
6
Overallestimated
57.0
Rank
#94
Percentile
Unavailable
Benchmarks
6

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

openaibenchlm:gpt-5-4-mini
primary
Input / 1M
$0.75
Cached input / 1M
$0.07
Output / 1M
$4.50
Context
400,000
View price source
Context window
400,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-03-17
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
37.46#120Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
46.58#102Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
57.30#45Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
44.60Not rankedNot publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
51.30Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
56.96#94Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model