Claude 4 Sonnet vs
Grok 4

Read the published evidence, route context, and missing facts before making a local decision. This page does not name a universal winner.

Models in this comparison

Provider identity and evidence state stay balanced across the pair.

A
Model type
Proprietary
Evidence state
Supported evidence
Published context
200,000
VS
X
Model type
Proprietary
Evidence state
Supported evidence
Published context
128,000

Key implications

Limited shared-metric coverage. Each finding is tied to a published metric, price, context, or modality fact.

  1. On Overall, Grok 4 has a higher supported BenchLM score (59.84 vs 42.37).
  2. Context window: Claude 4 Sonnet has the larger published context window (200,000 tokens vs 128,000 tokens).

Shared metric view

A radar is shown only when at least four compatible supported score metrics are published.

Comparable metric detail

  • AgenticClaude 4 Sonnet: 42.85 · Grok 4: Unavailable
  • CodingClaude 4 Sonnet: 43.69 · Grok 4: Unavailable
  • MathClaude 4 Sonnet: Unavailable · Grok 4: 38.6
  • OverallClaude 4 Sonnet: 42.37 · Grok 4: 59.84

Source metrics

Friendly metric names and published units stay visible. Missing measurements remain unavailable rather than becoming a score.

Source metric comparison
MetricUnitClaude 4 SonnetGrok 4
Agenticscore42.85Unavailable
Codingscore43.69Unavailable
MathscoreUnavailable38.6
Overallscore42.3759.84

Agentic

Unit
score
Claude 4 Sonnet
42.85
Grok 4
Unavailable

Coding

Unit
score
Claude 4 Sonnet
43.69
Grok 4
Unavailable

Math

Unit
score
Claude 4 Sonnet
Unavailable
Grok 4
38.6

Overall

Unit
score
Claude 4 Sonnet
42.37
Grok 4
59.84

Pricing and context

Verification is shown beside each selected route. Missing facts remain Not verified.

Route pricing and context comparison
FieldUnitClaude 4 SonnetGrok 4
Input API priceUSD / 1M tokens$3Not verified
Output API priceUSD / 1M tokens$15Not verified
Route contexttokens200,000Not verified
Input modalitiespublished listNot verifiedNot verified
Output modalitiespublished listNot verifiedNot verified

Input API price

Unit
USD / 1M tokens
Claude 4 Sonnet
$3
Grok 4
Not verified

Output API price

Unit
USD / 1M tokens
Claude 4 Sonnet
$15
Grok 4
Not verified

Route context

Unit
tokens
Claude 4 Sonnet
200,000
Grok 4
Not verified

Input modalities

Unit
published list
Claude 4 Sonnet
Not verified
Grok 4
Not verified

Output modalities

Unit
published list
Claude 4 Sonnet
Not verified
Grok 4
Not verified

Evidence provenance

Source records, route identity, timestamps, and methodology are consolidated here without declaring either model a winner.

Publication time
Aug 30, 2026, 2:15 AM UTC
Freshness
Stale — Published weekly benchmark evidence has not refreshed within 8 days.
Methodology
benchlm: benchlm_raw_composite
Model records
  • Claude 4 Sonnet — source benchlm · artifact models · model claude-4-sonnet
  • Grok 4 — source benchlm · artifact models · model grok-4
Selected price routes
  • Claude 4 Sonnet — route benchlm:claude-4-sonnet · source benchlm · provider anthropic
  • Grok 4 — Not published

Switch model pair

Choose from this result’s current and reviewed related models. Switching opens a reviewed comparison; it does not change this result’s evidence.

Step 1

Start with popular models, or search the full selectable directory.

Step 2

Start with popular models, or search the full selectable directory.