Battle result
5 configurations · 28 battles
quality · 2026-10-07T22:20:30.401Z · Final combined score
Recovered original evaluation; this was not a new model run. Equal transport and native system conditions are unverified.
Applied weights: quality 90%, tests 0%, speed 5%, cost 5%.
| Result | Rank 1Composer 2.5 · N/A | Rank 2Grok 4.7 · High | Rank 3DeepSeek Flash · N/A | Rank 4GPT-6 Luna · Low | Rank 5GPT-6.1 Sol · Low |
|---|---|---|---|---|---|
| Combined score | 80.1263 | 79.0105 | 72.9346 | 71.2032 | 70.3747 |
| Effort | N/A | High | N/A | Low | Low |
| Quality points | 75.6 | 77.4 | 66.6 | 63 | 66.6 |
| Test points | 0 | 0 | 0 | 0 | 0 |
| Speed points | 3.1414 | 1.2714 | 3.9064 | 4.2706 | 3.0404 |
| Cost points | 1.3849 | 0.3392 | 2.4282 | 3.9326 | 0.7343 |
| Time | 35.5 s | 175.963 s | 16.798 s | 10.247 s | 38.67 s |
| Reference USD | $0.0261025 | $0.137422 | $0.0105912 | $0.0027143 | $0.058092 |
Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.
Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.
Pairwise results
- Composer 2.5 · N/A vs DeepSeek Flash · N/A: First configuration wins
- Composer 2.5 · N/A vs GPT-6.1 Sol · Low: First configuration wins
- Composer 2.5 · N/A vs GPT-6 Luna · Low: First configuration wins
- Composer 2.5 · N/A vs Grok 4.7 · High: First configuration wins
- DeepSeek Flash · N/A vs GPT-6.1 Sol · Low: First configuration wins
- DeepSeek Flash · N/A vs GPT-6 Luna · Low: First configuration wins
- DeepSeek Flash · N/A vs Grok 4.7 · High: Second configuration wins
- GPT-6.1 Sol · Low vs GPT-6 Luna · Low: Second configuration wins
- GPT-6.1 Sol · Low vs Grok 4.7 · High: Second configuration wins
- GPT-6 Luna · Low vs Grok 4.7 · High: Second configuration wins