Battle result
5 configurations · 28 battles
quality · 2026-10-07T05:37:58.718Z · Final combined score
Recovered original evaluation; this was not a new model run. Equal transport and native system conditions are unverified.
Applied weights: quality 90%, tests 0%, speed 5%, cost 5%.
| Result | Rank 1GPT-6.1 Sol · Low | Rank 2Grok 4.7 · High | Rank 3GPT-6 Luna · Low | Rank 4DeepSeek Flash · N/A | Rank 5Composer 2.5 · N/A |
|---|---|---|---|---|---|
| Combined score | 84.9552 | 84.9102 | 80.4104 | 77.456 | 71.8975 |
| Effort | Low | High | Low | N/A | N/A |
| Quality points | 81 | 83.25 | 72 | 70.65 | 66.6 |
| Test points | 0 | 0 | 0 | 0 | 0 |
| Speed points | 2.5139 | 1.1191 | 3.8403 | 3.5814 | 3.2877 |
| Cost points | 1.4413 | 0.541 | 4.57 | 3.2246 | 2.0098 |
| Time | 59.336 s | 208.071 s | 18.118 s | 23.766 s | 31.248 s |
| Reference USD | $0.024692 | $0.082414 | $0.0009408 | $0.0055059 | $0.0148785 |
Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.
Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.
Pairwise results
- DeepSeek Flash · N/A vs Composer 2.5 · N/A: First configuration wins
- DeepSeek Flash · N/A vs Grok 4.7 · High: Second configuration wins
- DeepSeek Flash · N/A vs GPT-6.1 Sol · Low: Second configuration wins
- DeepSeek Flash · N/A vs GPT-6 Luna · Low: Second configuration wins
- Composer 2.5 · N/A vs Grok 4.7 · High: Second configuration wins
- Composer 2.5 · N/A vs GPT-6.1 Sol · Low: Second configuration wins
- Composer 2.5 · N/A vs GPT-6 Luna · Low: Second configuration wins
- Grok 4.7 · High vs GPT-6.1 Sol · Low: Second configuration wins
- Grok 4.7 · High vs GPT-6 Luna · Low: First configuration wins
- GPT-6.1 Sol · Low vs GPT-6 Luna · Low: First configuration wins