Key implications
Insufficient shared-metric coverage. Each finding is tied to a published metric, price, context, or modality fact.
- Input API price: Grok 4.20 has the lower verified rate ($2 / 1M tokens vs $5 / 1M tokens).
- Output API price: Grok 4.20 has the lower verified rate ($6 / 1M tokens vs $25 / 1M tokens).
- Context window: Grok 4.20 has the larger published context window (2,000,000 tokens vs 1,000,000 tokens).
Shared metric view
A radar is shown only when at least four compatible supported score metrics are published.
Comparable metric detail
- AgenticClaude Opus 4.7: 58.5 · Grok 4.20: 50.89
- CodingClaude Opus 4.7: 68.72 · Grok 4.20: 46.12
- KnowledgeClaude Opus 4.7: 63.9 · Grok 4.20: Unavailable
- MathClaude Opus 4.7: 61.8 · Grok 4.20: Unavailable
- MultimodalGroundedClaude Opus 4.7: Unavailable · Grok 4.20: 30.5
- ReasoningClaude Opus 4.7: Unavailable · Grok 4.20: 52.7
- OverallClaude Opus 4.7: 72.33 · Grok 4.20: 55.44
Source metrics
Friendly metric names and published units stay visible. Missing measurements remain unavailable rather than becoming a score.
Source metric comparison| Metric | Unit | Claude Opus 4.7 | Grok 4.20 |
|---|
| Agentic | score | 58.5 | 50.89 |
|---|
| Coding | score | 68.72 | 46.12 |
|---|
| Knowledge | score | 63.9 | Unavailable |
|---|
| Math | score | 61.8 | Unavailable |
|---|
| MultimodalGrounded | score | Unavailable | 30.5 |
|---|
| Reasoning | score | Unavailable | 52.7 |
|---|
| Overall | score | 72.33 | 55.44 |
|---|
Agentic
- Unit
- score
- Claude Opus 4.7
- 58.5
- Grok 4.20
- 50.89
Coding
- Unit
- score
- Claude Opus 4.7
- 68.72
- Grok 4.20
- 46.12
Knowledge
- Unit
- score
- Claude Opus 4.7
- 63.9
- Grok 4.20
- Unavailable
Math
- Unit
- score
- Claude Opus 4.7
- 61.8
- Grok 4.20
- Unavailable
MultimodalGrounded
- Unit
- score
- Claude Opus 4.7
- Unavailable
- Grok 4.20
- 30.5
Reasoning
- Unit
- score
- Claude Opus 4.7
- Unavailable
- Grok 4.20
- 52.7
Overall
- Unit
- score
- Claude Opus 4.7
- 72.33
- Grok 4.20
- 55.44
Pricing and context
Verification is shown beside each selected route. Missing facts remain Not verified.
Route pricing and context comparison| Field | Unit | Claude Opus 4.7 | Grok 4.20 |
|---|
| Input API price | USD / 1M tokens | $5 | $2 |
|---|
| Output API price | USD / 1M tokens | $25 | $6 |
|---|
| Route context | tokens | 1,000,000 | 2,000,000 |
|---|
| Input modalities | published list | Not verified | Not verified |
|---|
| Output modalities | published list | Not verified | Not verified |
|---|
Input API price
- Unit
- USD / 1M tokens
- Claude Opus 4.7
- $5
- Grok 4.20
- $2
Output API price
- Unit
- USD / 1M tokens
- Claude Opus 4.7
- $25
- Grok 4.20
- $6
Route context
- Unit
- tokens
- Claude Opus 4.7
- 1,000,000
- Grok 4.20
- 2,000,000
Input modalities
- Unit
- published list
- Claude Opus 4.7
- Not verified
- Grok 4.20
- Not verified
Output modalities
- Unit
- published list
- Claude Opus 4.7
- Not verified
- Grok 4.20
- Not verified
Evidence provenance
Source records, route identity, timestamps, and methodology are consolidated here without declaring either model a winner.
- Publication time
- Aug 30, 2026, 2:15 AM UTC
- Freshness
- Stale — Published weekly benchmark evidence has not refreshed within 8 days.
- Methodology
- benchlm: benchlm_raw_composite
- Model records
- Claude Opus 4.7 — source benchlm · artifact models · model claude-opus-4-7
- Grok 4.20 — source benchlm · artifact models · model grok-4-20-beta
- Selected price routes
- Claude Opus 4.7 — route benchlm:claude-opus-4-7 · source benchlm · provider anthropic
- Grok 4.20 — route benchlm:grok-4-20-beta · source benchlm · provider xai
Switch model pair
Choose from this result’s current and reviewed related models. Switching opens a reviewed comparison; it does not change this result’s evidence.