Model evidence profile

Claude Sonnet 4.5

Anthropic · Proprietary · Non-Reasoning

currentestimated
Overall public score
54.49
Source rank
#108
Evidence coverage
6 benchmarks · 2 sources

Strongest published evidencePublic overall score 54.49 at source rank #108.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMathMultimodal groundedOverallReasoning
Agentic percentile
43.9 percentileRank #79 of 140
Coding percentile
52.8 percentileRank #69 of 145
Knowledge percentile
Knowledge: Unavailable
Math percentile
Math: Unavailable
Multimodal grounded percentile
Multimodal grounded: Unavailable
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
46.7
Rank
#79 of 140
Percentile
43.9%
Benchmarks
6
Codingestimated
51.8
Rank
#69 of 145
Percentile
52.8%
Benchmarks
6
Knowledgeestimated
74.7
Rank
Not ranked
Percentile
Unavailable
Benchmarks
6
Mathestimated
34.9
Rank
Not ranked
Percentile
Unavailable
Benchmarks
6
Overallestimated
54.5
Rank
#108
Percentile
Unavailable
Benchmarks
6
Reasoningestimated
19.7
Rank
Not ranked
Percentile
Unavailable
Benchmarks
6

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

anthropicbenchlm:claude-sonnet-4-5
primary
Input / 1M
$3.00
Cached input / 1M
Unavailable
Output / 1M
$15.00
Context
200,000
View price source
Context window
200,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2025-09-01
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
46.70#79Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
51.82#69Not publishedAug 29, 2026BenchLM

knowledge

ScoreRankWeightLast UpdatedSource
74.70Not rankedNot publishedAug 29, 2026BenchLM

math

ScoreRankWeightLast UpdatedSource
34.90Not rankedNot publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
19.70Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
54.49#108Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model