GPT-5.6 Sol vs
Muse Spark 1.1

Read the published evidence, route context, and missing facts before making a local decision. This page does not name a universal winner.

Models in this comparison

Provider identity and evidence state stay balanced across the pair.

Model type
Proprietary
Evidence state
Supported evidence
Published context
1,050,000
VS
Model type
Proprietary
Evidence state
Supported evidence
Published context
1,000,000

Key implications

Broad shared-metric coverage. Each finding is tied to a published metric, price, context, or modality fact.

  1. Across compatible supported BenchLM categories, GPT-5.6 Sol has higher scores in Coding, Agentic, and Overall; Muse Spark 1.1 has a higher score in Knowledge.
  2. Context window: GPT-5.6 Sol has the larger published context window (1,050,000 tokens vs 1,000,000 tokens).

Shared metric view

The radar is a per-axis relative view; exact values and units remain available in its adjacent table.

Per-axis relative view: each axis scales to the higher published value for that exact shared metric. The table preserves the exact published values and units.
  • GPT-5.6 Sol: solid line
  • Muse Spark 1.1: dashed line
GPT-5.6 Sol and Muse Spark 1.1 shared metric radarAgenticCodingKnowledgeOverall
Exact published values used in the per-axis relative radar chart.
MetricGPT-5.6 SolMuse Spark 1.1Unit
Agentic68.8859.82score
Coding79.0165.7score
Knowledge8692.8score
Overall82.277.14score

Source metrics

Friendly metric names and published units stay visible. Missing measurements remain unavailable rather than becoming a score.

Source metric comparison
MetricUnitGPT-5.6 SolMuse Spark 1.1
Agenticscore68.8859.82
Codingscore79.0165.7
Knowledgescore8692.8
Mathscore97Unavailable
MultimodalGroundedscore82.772
Reasoningscore85.2Unavailable
Overallscore82.277.14

Agentic

Unit
score
GPT-5.6 Sol
68.88
Muse Spark 1.1
59.82

Coding

Unit
score
GPT-5.6 Sol
79.01
Muse Spark 1.1
65.7

Knowledge

Unit
score
GPT-5.6 Sol
86
Muse Spark 1.1
92.8

Math

Unit
score
GPT-5.6 Sol
97
Muse Spark 1.1
Unavailable

MultimodalGrounded

Unit
score
GPT-5.6 Sol
82.7
Muse Spark 1.1
72

Reasoning

Unit
score
GPT-5.6 Sol
85.2
Muse Spark 1.1
Unavailable

Overall

Unit
score
GPT-5.6 Sol
82.2
Muse Spark 1.1
77.14

Pricing and context

Verification is shown beside each selected route. Missing facts remain Not verified.

Route pricing and context comparison
FieldUnitGPT-5.6 SolMuse Spark 1.1
Input API priceUSD / 1M tokens$5Not verified
Cached input API priceUSD / 1M tokens$0.5Not verified
Output API priceUSD / 1M tokens$30Not verified
Route contexttokens1,050,000Not verified
Input modalitiespublished listNot verifiedNot verified
Output modalitiespublished listNot verifiedNot verified

Input API price

Unit
USD / 1M tokens
GPT-5.6 Sol
$5
Muse Spark 1.1
Not verified

Cached input API price

Unit
USD / 1M tokens
GPT-5.6 Sol
$0.5
Muse Spark 1.1
Not verified

Output API price

Unit
USD / 1M tokens
GPT-5.6 Sol
$30
Muse Spark 1.1
Not verified

Route context

Unit
tokens
GPT-5.6 Sol
1,050,000
Muse Spark 1.1
Not verified

Input modalities

Unit
published list
GPT-5.6 Sol
Not verified
Muse Spark 1.1
Not verified

Output modalities

Unit
published list
GPT-5.6 Sol
Not verified
Muse Spark 1.1
Not verified

Evidence provenance

Source records, route identity, timestamps, and methodology are consolidated here without declaring either model a winner.

Publication time
Aug 30, 2026, 2:15 AM UTC
Freshness
Stale — Published weekly benchmark evidence has not refreshed within 8 days.
Methodology
benchlm: benchlm_raw_composite
Model records
  • GPT-5.6 Sol — source benchlm · artifact models · model gpt-5-6-sol
  • Muse Spark 1.1 — source benchlm · artifact models · model muse-spark-1-1
Selected price routes
  • GPT-5.6 Sol — route benchlm:gpt-5-6-sol · source benchlm · provider openai
  • Muse Spark 1.1 — Not published

Other reviewed matchups from the same published revision, ready to open without changing this result’s evidence.

Switch model pair

Choose from this result’s current and reviewed related models. Switching opens a reviewed comparison; it does not change this result’s evidence.

Step 1

Start with popular models, or search the full selectable directory.

Step 2

Start with popular models, or search the full selectable directory.