Model evidence profile

GPT-5.4

OpenAI · Proprietary · Reasoning

currentsupported
Overall public score
73.56
Source rank
#12
Evidence coverage
7 benchmarks · 2 sources

Strongest published evidencePublic overall score 73.56 at source rank #12.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
79.1 percentileRank #30 of 140
Coding percentile
74.3 percentileRank #38 of 145
Knowledge percentile
64.8 percentileRank #20 of 55
Math percentile
Math: Unavailable
Multimodal grounded percentile
55.9 percentileRank #16 of 35
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticsupported
57.5
Rank
#30 of 140
Percentile
79.1%
Benchmarks
7
Codingsupported
58.9
Rank
#38 of 145
Percentile
74.3%
Benchmarks
7
Knowledgesupported
75.3
Rank
#20 of 55
Percentile
64.8%
Benchmarks
7
Mathsupported
65.7
Rank
Not ranked
Percentile
Unavailable
Benchmarks
7
Overallsupported
73.6
Rank
#12
Percentile
Unavailable
Benchmarks
7
Reasoningsupported
69.8
Rank
Not ranked
Percentile
Unavailable
Benchmarks
7

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

openaibenchlm:gpt-5-4
primary
Input / 1M
$2.50
Cached input / 1M
$0.25
Output / 1M
$15.00
Context
1,050,000
View price source
Context window
1,050,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-03-05
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
57.52#30Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
58.90#38Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
75.30#20Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
65.70Not rankedNot publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
63.50#16Not publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
69.80Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
73.56#12Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model