Winner

Grok 4.7 Low

Engineering · October 8, 2026 · Difficulty: Medium

Final combined results in supplied rank order. Tied ranks remain tied.
ResultRank 1Grok 4.7 LowRank 2GPT-6.1 Sol LowRank 3Composer 2.5 DefaultRank 4GPT-6 Luna LowRank 5Grok 4.7 HighRank 6GPT-6.1 Sol HighRank 7GPT-6 Luna Extra HighRank 8DeepSeek Flash N/A
Combined score85.6080.5976.0172.8172.6072.4761.5653.83
Judge rank12364578
Judge score /10090.0082.0077.0072.0076.5074.5062.0052.00
Time882.24 s48.56 s43.27 s8.68 s2,392.58 s106.61 s51.67 s9.18 s
Estimated USDN/AN/AN/AN/AN/AN/AN/AN/A

Estimated reference costs are not billed charges. Judge rank is derived from recorded rubric metrics across judged configurations; final rank uses the combined score.

Component points
Final combined results in supplied rank order. Tied ranks remain tied.
ResultRank 1 Grok 4.7 LowRank 2 GPT-6.1 Sol LowRank 3 Composer 2.5 DefaultRank 4 GPT-6 Luna LowRank 5 Grok 4.7 HighRank 6 GPT-6.1 Sol HighRank 7 GPT-6 Luna Extra HighRank 8 DeepSeek Flash N/A
Quality points85.263277.684272.947468.210572.473770.578958.736849.2632
Test points00000000
Speed points0.33512.90893.05784.59830.12881.89542.82794.5648
Cost points00000000

Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.

Advanced results and scoring details

engineering · 2026-10-08T21:38:24.033Z · Prospective evaluation. Equal transport and native system conditions are unverified.

Applied weights: quality 94.7368%, tests 0%, speed 5.2632%, cost 0%.

Cost was omitted for every configuration in this cohort. Active weights were normalized uniformly. Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.

Pairwise results

  • GPT-6.1 Sol High vs DeepSeek Flash N/A: First configuration wins
  • GPT-6.1 Sol High vs GPT-6.1 Sol Low: Second configuration wins
  • GPT-6.1 Sol High vs GPT-6 Luna Low: Second configuration wins
  • GPT-6.1 Sol High vs GPT-6 Luna Extra High: First configuration wins
  • GPT-6.1 Sol High vs Composer 2.5 Default: Second configuration wins
  • GPT-6.1 Sol High vs Grok 4.7 Low: Second configuration wins
  • GPT-6.1 Sol High vs Grok 4.7 High: Second configuration wins
  • DeepSeek Flash N/A vs GPT-6.1 Sol Low: Second configuration wins
  • DeepSeek Flash N/A vs GPT-6 Luna Low: Second configuration wins
  • DeepSeek Flash N/A vs GPT-6 Luna Extra High: Second configuration wins
  • DeepSeek Flash N/A vs Composer 2.5 Default: Second configuration wins
  • DeepSeek Flash N/A vs Grok 4.7 Low: Second configuration wins
  • DeepSeek Flash N/A vs Grok 4.7 High: Second configuration wins
  • GPT-6.1 Sol Low vs GPT-6 Luna Low: First configuration wins
  • GPT-6.1 Sol Low vs GPT-6 Luna Extra High: First configuration wins
  • GPT-6.1 Sol Low vs Composer 2.5 Default: First configuration wins
  • GPT-6.1 Sol Low vs Grok 4.7 Low: Second configuration wins
  • GPT-6.1 Sol Low vs Grok 4.7 High: First configuration wins
  • GPT-6 Luna Low vs GPT-6 Luna Extra High: First configuration wins
  • GPT-6 Luna Low vs Composer 2.5 Default: Second configuration wins
  • GPT-6 Luna Low vs Grok 4.7 Low: Second configuration wins
  • GPT-6 Luna Low vs Grok 4.7 High: First configuration wins
  • GPT-6 Luna Extra High vs Composer 2.5 Default: Second configuration wins
  • GPT-6 Luna Extra High vs Grok 4.7 Low: Second configuration wins
  • GPT-6 Luna Extra High vs Grok 4.7 High: Second configuration wins
  • Composer 2.5 Default vs Grok 4.7 Low: Second configuration wins
  • Composer 2.5 Default vs Grok 4.7 High: First configuration wins
  • Grok 4.7 Low vs Grok 4.7 High: First configuration wins