Model evidence profile

GPT-4.1 mini

OpenAI · Proprietary · Non-Reasoning

currentestimated
Overall public score
44.33
Source rank
#174
Evidence coverage
5 benchmarks · 2 sources

Strongest published evidencePublic overall score 44.33 at source rank #174.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingInstruction FollowingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
25.9 percentileRank #104 of 140
Coding percentile
25.7 percentileRank #108 of 145
Instruction Following percentile
34.1 percentileRank #28 of 42
Knowledge percentile
Knowledge: Unavailable
Math percentile
Math: Unavailable
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
41.7
Rank
#104 of 140
Percentile
25.9%
Benchmarks
5
Codingestimated
45.0
Rank
#108 of 145
Percentile
25.7%
Benchmarks
5
Instruction Followingestimated
75.5
Rank
#28 of 42
Percentile
34.1%
Benchmarks
5
Knowledgeestimated
55.4
Rank
Not ranked
Percentile
Unavailable
Benchmarks
5
Mathestimated
28.6
Rank
Not ranked
Percentile
Unavailable
Benchmarks
5
Overallestimated
44.3
Rank
#174
Percentile
Unavailable
Benchmarks
5

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

openaibenchlm:gpt-4-1-mini
primary
Input / 1M
$0.40
Cached input / 1M
Unavailable
Output / 1M
$1.60
Context
1,000,000
View price source
Context window
1,000,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2025-04-14
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
41.74#104Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
45.04#108Not publishedAug 29, 2026BenchLM

instructionFollowing

ScoreRankWeightLast UpdatedSource
75.50#28Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
55.40Not rankedNot publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
28.60Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
44.33#174Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model