Model evidence profile

Claude Opus 4.8

Anthropic · Proprietary · Reasoning

currentsupported
Overall public score
76.60
Source rank
#9
Evidence coverage
7 benchmarks · 2 sources

Strongest published evidencePublic overall score 76.6 at source rank #9.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
87.8 percentileRank #18 of 140
Coding percentile
95.1 percentileRank #8 of 145
Knowledge percentile
96.3 percentileRank #3 of 55
Math percentile
83.3 percentileRank #2 of 7
Multimodal grounded percentile
91.2 percentileRank #4 of 35
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticsupported
61.8
Rank
#18 of 140
Percentile
87.8%
Benchmarks
7
Codingsupported
72.0
Rank
#8 of 145
Percentile
95.1%
Benchmarks
7
Knowledgesupported
86.8
Rank
#3 of 55
Percentile
96.3%
Benchmarks
7
Mathsupported
66.8
Rank
#2 of 7
Percentile
83.3%
Benchmarks
7
Overallsupported
76.6
Rank
#9
Percentile
Unavailable
Benchmarks
7
Reasoningsupported
68.3
Rank
Not ranked
Percentile
Unavailable
Benchmarks
7

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

anthropicbenchlm:claude-opus-4-8
primary
Input / 1M
$5.00
Cached input / 1M
Unavailable
Output / 1M
$25.00
Context
1,000,000
View price source
Context window
1,000,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-05-28
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
61.78#18Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
71.98#8Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
86.80#3Not publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
66.80#2Not publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
85.40#4Not publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
68.30Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
76.60#9Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model