Winner

GPT-6.1 Sol Max

Orchestrators · October 8, 2026 · Difficulty: High

Final combined results in supplied rank order. Tied ranks remain tied.
ResultRank 1GPT-6.1 Sol MaxRank 2GPT-6 Astra LowRank 3GPT-6 Astra HighRank 4GPT-6.1 Sol LowRank 5GPT-6 Luna LowRank 6Composer 2.5 Default
Combined score88.0485.8480.8979.7241.5035.54
Judge rank123456
Judge score /10092.5087.5084.0081.0039.0035.00
Judge breakdown
Judge breakdown
Correctness /10
9
Coverage /10
10
Evidence /10
9
Actionability /10
9
Judge breakdown
Correctness /10
9
Coverage /10
8
Evidence /10
9
Actionability /10
9
Judge breakdown
Correctness /10
9
Coverage /10
8
Evidence /10
8
Actionability /10
8
Judge breakdown
Correctness /10
8
Coverage /10
7
Evidence /10
9
Actionability /10
9
Judge breakdown
Correctness /10
3
Coverage /10
6
Evidence /10
3
Actionability /10
4
Judge breakdown
Correctness /10
2
Coverage /10
5
Evidence /10
5
Actionability /10
3
Time717.4 s47.28 s181.28 s45.95 s9.3 s72.57 s
Estimated USD$0.451714N/AN/A$0.194894$0.0092994$0.071348

Reference estimates use original frozen rates; cost was excluded from this battle score; not billed USD.

Judge rank is derived from recorded rubric metrics across judged configurations; final rank uses the combined score.

Component points
Final combined results in supplied rank order. Tied ranks remain tied.
ResultRank 1 GPT-6.1 Sol MaxRank 2 GPT-6 Astra LowRank 3 GPT-6 Astra HighRank 4 GPT-6.1 Sol LowRank 5 GPT-6 Luna LowRank 6 Composer 2.5 Default
Quality points87.631682.894779.578976.736836.947433.1579
Test points000000
Speed points0.40622.94371.30882.98064.55692.3821
Cost points000000

Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.

Advanced results and scoring details

orchestrators · 2026-10-08T23:46:33.692Z · Prospective evaluation. Equal transport and native system conditions are unverified.

Applied weights: quality 94.7368%, tests 0%, speed 5.2632%, cost 0%.

Cost was omitted for every configuration in this cohort. Active weights were normalized uniformly. Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.

Pairwise results

  • Composer 2.5 Default vs GPT-6 Astra High: Second configuration wins
  • Composer 2.5 Default vs GPT-6 Luna Low: Second configuration wins
  • Composer 2.5 Default vs GPT-6.1 Sol Low: Second configuration wins
  • Composer 2.5 Default vs GPT-6 Astra Low: Second configuration wins
  • Composer 2.5 Default vs GPT-6.1 Sol Max: Second configuration wins
  • GPT-6 Astra High vs GPT-6 Luna Low: First configuration wins
  • GPT-6 Astra High vs GPT-6.1 Sol Low: First configuration wins
  • GPT-6 Astra High vs GPT-6 Astra Low: Second configuration wins
  • GPT-6 Astra High vs GPT-6.1 Sol Max: Second configuration wins
  • GPT-6 Luna Low vs GPT-6.1 Sol Low: Second configuration wins
  • GPT-6 Luna Low vs GPT-6 Astra Low: Second configuration wins
  • GPT-6 Luna Low vs GPT-6.1 Sol Max: Second configuration wins
  • GPT-6.1 Sol Low vs GPT-6 Astra Low: Second configuration wins
  • GPT-6.1 Sol Low vs GPT-6.1 Sol Max: Second configuration wins
  • GPT-6 Astra Low vs GPT-6.1 Sol Max: Second configuration wins