Model evidence profile

Grok 4.20

xAI · Proprietary · Reasoning

currentestimated
Overall public score
55.44
Source rank
#101
Evidence coverage
5 benchmarks · 2 sources

Strongest published evidencePublic overall score 55.44 at source rank #101.

Validate before choosingValidate the selected route price, context limits, and evidence freshness before choosing.

Revision benchmark_3b14dc55ef5307cfe91356d3ea348101 · Published Aug 30, 2026 · Checked Aug 30, 2026 · stale

Relative field position

Capability radar

Percentiles use eligible source ranks. Missing axes remain unavailable and never become zero.

Capability ranking percentile radarAgenticCodingKnowledgeMultimodal groundedOverallReasoning
Agentic percentile
57.6 percentileRank #60 of 140
Coding percentile
27.8 percentileRank #105 of 145
Knowledge percentile
Knowledge: Unavailable
Multimodal grounded percentile
11.8 percentileRank #31 of 35
Overall percentile
Overall: Unavailable
Reasoning percentile
Reasoning: Unavailable

Published measurements

Category scores

Scores retain their source units; percentile and rank appear only for eligible fields.

Agenticestimated
50.9
Rank
#60 of 140
Percentile
57.6%
Benchmarks
5
Codingestimated
46.1
Rank
#105 of 145
Percentile
27.8%
Benchmarks
5
Overallestimated
55.4
Rank
#101
Percentile
Unavailable
Benchmarks
5
Reasoningestimated
52.7
Rank
Not ranked
Percentile
Unavailable
Benchmarks
5

Route-specific facts

Pricing and specifications

Conflicting routes remain separate and attributable.

xaibenchlm:grok-4-20-beta
primary
Input / 1M
$2.00
Cached input / 1M
Unavailable
Output / 1M
$6.00
Context
2,000,000
View price source
Context window
2,000,000
Maximum output
Unavailable
Input modalities
Unavailable
Output modalities
Unavailable
Release date
2026-03-10
Self hosting
Not verified

Auditable evidence

Benchmark ledger

Display values, source ranks, and provenance remain visible without implying unsupported aggregate weight.

agentic

ScoreRankWeightLast UpdatedSource
50.89#60Not publishedAug 29, 2026BenchLM

coding

ScoreRankWeightLast UpdatedSource
46.12#105Not publishedAug 29, 2026BenchLM

multimodalGrounded

ScoreRankWeightLast UpdatedSource
30.50#31Not publishedAug 29, 2026BenchLM

reasoning

ScoreRankWeightLast UpdatedSource
52.70Not rankedNot publishedAug 29, 2026BenchLM

overall

ScoreRankWeightLast UpdatedSource
55.44#101Not publishedAug 29, 2026BenchLM

Related evidence

Compare this model