Battle result
5 configurations · 28 battles
quality · 2026-10-07T07:03:00.538Z · Final combined score
Recovered original evaluation; this was not a new model run. Equal transport and native system conditions are unverified.
Applied weights: quality 90%, tests 0%, speed 5%, cost 5%.
| Result | Rank 1GPT-6.1 Sol · Low | Rank 2DeepSeek Flash · N/A | Rank 3GPT-6 Luna · Low | Rank 4Grok 4.7 · High | Rank 5Composer 2.5 · N/A |
|---|---|---|---|---|---|
| Combined score | 86.8936 | 81.6151 | 80.1636 | 75.7372 | 69.2465 |
| Effort | Low | N/A | Low | High | N/A |
| Quality points | 83.25 | 74.25 | 72 | 74.25 | 63.9 |
| Test points | 0 | 0 | 0 | 0 | 0 |
| Speed points | 2.2089 | 3.8246 | 3.5558 | 0.951 | 3.2728 |
| Cost points | 1.4346 | 3.5405 | 4.6077 | 0.5362 | 2.0738 |
| Time | 75.813 s | 18.44 s | 24.369 s | 255.463 s | 31.666 s |
| Reference USD | $0.024852 | $0.0041223 | $0.0008513 | $0.08325 | $0.0141105 |
Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.
Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.
Pairwise results
- GPT-6.1 Sol · Low vs GPT-6 Luna · Low: First configuration wins
- GPT-6.1 Sol · Low vs Composer 2.5 · N/A: First configuration wins
- GPT-6.1 Sol · Low vs Grok 4.7 · High: First configuration wins
- GPT-6.1 Sol · Low vs DeepSeek Flash · N/A: First configuration wins
- GPT-6 Luna · Low vs Composer 2.5 · N/A: First configuration wins
- GPT-6 Luna · Low vs Grok 4.7 · High: First configuration wins
- GPT-6 Luna · Low vs DeepSeek Flash · N/A: Second configuration wins
- Composer 2.5 · N/A vs Grok 4.7 · High: Second configuration wins
- Composer 2.5 · N/A vs DeepSeek Flash · N/A: Second configuration wins
- Grok 4.7 · High vs DeepSeek Flash · N/A: Second configuration wins