Model evidence profile

o3-mini

OpenAI · Proprietary · Reasoning

currentsupported
Overall public score
46.94
Source rank
#159
Evidence coverage
4 benchmarks · 2 sources

Strongest published evidencePublic overall score 46.94 at source rank #159.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingInstruction FollowingKnowledgeMultimodal groundedOverallReasoning
Agentic percentile
Agentic: Unavailable
Coding percentile
36.8 percentileRank #92 of 145
Instruction Following percentile
73.2 percentileRank #12 of 42
Knowledge percentile
Knowledge: Unavailable
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Codingsupported
47.5
Rank
#92 of 145
Percentile
36.8%
Benchmarks
4
Instruction Followingsupported
91.2
Rank
#12 of 42
Percentile
73.2%
Benchmarks
4
Knowledgesupported
68.5
Rank
Not ranked
Percentile
Unavailable
Benchmarks
4
Overallsupported
46.9
Rank
#159
Percentile
Unavailable
Benchmarks
4

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

openaibenchlm:o3-mini
primary
Input / 1M
$1.10
Cached input / 1M
Unavailable
Output / 1M
$4.40
Context
200,000
View price source
Context window
200,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2025-01-31
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

coding

ScoreRankWeightLast UpdatedSource
47.54#92Not publishedAug 29, 2026BenchLM

instructionFollowing

ScoreRankWeightLast UpdatedSource
91.20#12Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
68.50Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
46.94#159Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model