| Result | Rank 1Grok 4.7 Low | Rank 2GPT-6.1 Sol Low | Rank 3Composer 2.5 Default | Rank 4GPT-6 Luna Low | Rank 5Grok 4.7 High | Rank 6GPT-6.1 Sol High | Rank 7GPT-6 Luna Extra High | Rank 8DeepSeek Flash N/A |
|---|---|---|---|---|---|---|---|---|
| Combined score | 85.60 | 80.59 | 76.01 | 72.81 | 72.60 | 72.47 | 61.56 | 53.83 |
| Judge rank | 1 | 2 | 3 | 6 | 4 | 5 | 7 | 8 |
| Judge score /100 | 90.00 | 82.00 | 77.00 | 72.00 | 76.50 | 74.50 | 62.00 | 52.00 |
| Time | 882.24 s | 48.56 s | 43.27 s | 8.68 s | 2,392.58 s | 106.61 s | 51.67 s | 9.18 s |
| Estimated USD | N/A | N/A | N/A | N/A | N/A | N/A | N/A | N/A |
Estimated reference costs are not billed charges. Judge rank is derived from recorded rubric metrics across judged configurations; final rank uses the combined score.
Component points
| Result | Rank 1 Grok 4.7 Low | Rank 2 GPT-6.1 Sol Low | Rank 3 Composer 2.5 Default | Rank 4 GPT-6 Luna Low | Rank 5 Grok 4.7 High | Rank 6 GPT-6.1 Sol High | Rank 7 GPT-6 Luna Extra High | Rank 8 DeepSeek Flash N/A |
|---|---|---|---|---|---|---|---|---|
| Quality points | 85.2632 | 77.6842 | 72.9474 | 68.2105 | 72.4737 | 70.5789 | 58.7368 | 49.2632 |
| Test points | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| Speed points | 0.3351 | 2.9089 | 3.0578 | 4.5983 | 0.1288 | 1.8954 | 2.8279 | 4.5648 |
| Cost points | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.
Advanced results and scoring details
engineering · 2026-10-08T21:38:24.033Z · Prospective evaluation. Equal transport and native system conditions are unverified.
Applied weights: quality 94.7368%, tests 0%, speed 5.2632%, cost 0%.
Cost was omitted for every configuration in this cohort. Active weights were normalized uniformly. Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.
Pairwise results
- GPT-6.1 Sol High vs DeepSeek Flash N/A: First configuration wins
- GPT-6.1 Sol High vs GPT-6.1 Sol Low: Second configuration wins
- GPT-6.1 Sol High vs GPT-6 Luna Low: Second configuration wins
- GPT-6.1 Sol High vs GPT-6 Luna Extra High: First configuration wins
- GPT-6.1 Sol High vs Composer 2.5 Default: Second configuration wins
- GPT-6.1 Sol High vs Grok 4.7 Low: Second configuration wins
- GPT-6.1 Sol High vs Grok 4.7 High: Second configuration wins
- DeepSeek Flash N/A vs GPT-6.1 Sol Low: Second configuration wins
- DeepSeek Flash N/A vs GPT-6 Luna Low: Second configuration wins
- DeepSeek Flash N/A vs GPT-6 Luna Extra High: Second configuration wins
- DeepSeek Flash N/A vs Composer 2.5 Default: Second configuration wins
- DeepSeek Flash N/A vs Grok 4.7 Low: Second configuration wins
- DeepSeek Flash N/A vs Grok 4.7 High: Second configuration wins
- GPT-6.1 Sol Low vs GPT-6 Luna Low: First configuration wins
- GPT-6.1 Sol Low vs GPT-6 Luna Extra High: First configuration wins
- GPT-6.1 Sol Low vs Composer 2.5 Default: First configuration wins
- GPT-6.1 Sol Low vs Grok 4.7 Low: Second configuration wins
- GPT-6.1 Sol Low vs Grok 4.7 High: First configuration wins
- GPT-6 Luna Low vs GPT-6 Luna Extra High: First configuration wins
- GPT-6 Luna Low vs Composer 2.5 Default: Second configuration wins
- GPT-6 Luna Low vs Grok 4.7 Low: Second configuration wins
- GPT-6 Luna Low vs Grok 4.7 High: First configuration wins
- GPT-6 Luna Extra High vs Composer 2.5 Default: Second configuration wins
- GPT-6 Luna Extra High vs Grok 4.7 Low: Second configuration wins
- GPT-6 Luna Extra High vs Grok 4.7 High: Second configuration wins
- Composer 2.5 Default vs Grok 4.7 Low: Second configuration wins
- Composer 2.5 Default vs Grok 4.7 High: First configuration wins
- Grok 4.7 Low vs Grok 4.7 High: First configuration wins