Battle result
5 configurations · 28 battles
quality · 2026-10-07T22:17:38.500Z · Final combined score
Recovered original evaluation; this was not a new model run. Equal transport and native system conditions are unverified.
Applied weights: quality 90%, tests 0%, speed 5%, cost 5%.
| Result | Rank 1Composer 2.5 · N/A | Rank 2GPT-6.1 Sol · Low | Rank 3Grok 4.7 · High | Rank 4GPT-6 Luna · Low | Rank 5DeepSeek Flash · N/A |
|---|---|---|---|---|---|
| Combined score | 89.9811 | 77.5719 | 77.0036 | 74.1151 | 64.052 |
| Effort | N/A | Low | High | Low | N/A |
| Quality points | 84.6 | 72 | 75.6 | 64.8 | 55.8 |
| Test points | 0 | 0 | 0 | 0 | 0 |
| Speed points | 3.4364 | 3.5482 | 0.9923 | 4.5698 | 4.3634 |
| Cost points | 1.9447 | 2.0236 | 0.4113 | 4.7454 | 3.8886 |
| Time | 27.301 s | 24.549 s | 242.325 s | 5.649 s | 8.753 s |
| Reference USD | $0.0157105 | $0.014708 | $0.11158 | $0.0005366 | $0.0028581 |
Swipe, scroll or use Previous and More to compare configurations. Row labels stay visible.
Displayed scores are rounded; pairwise outcomes use exact recorded fractions. Reference costs are estimates, not billed charges. Missing components follow the recorded assessment.
Pairwise results
- Grok 4.7 · High vs GPT-6.1 Sol · Low: Second configuration wins
- Grok 4.7 · High vs DeepSeek Flash · N/A: First configuration wins
- Grok 4.7 · High vs GPT-6 Luna · Low: First configuration wins
- Grok 4.7 · High vs Composer 2.5 · N/A: Second configuration wins
- GPT-6.1 Sol · Low vs DeepSeek Flash · N/A: First configuration wins
- GPT-6.1 Sol · Low vs GPT-6 Luna · Low: First configuration wins
- GPT-6.1 Sol · Low vs Composer 2.5 · N/A: Second configuration wins
- DeepSeek Flash · N/A vs GPT-6 Luna · Low: Second configuration wins
- DeepSeek Flash · N/A vs Composer 2.5 · N/A: Second configuration wins
- GPT-6 Luna · Low vs Composer 2.5 · N/A: Second configuration wins