Published battle record

Quality battle

5 models share the highest score.

Showing 3 of 10 models. Full scorecards below.

QualityQuality battle

Shared victory

. 10 participants. Difficulty: Not recorded. Quality decides the result.

How the score became the result

The quality bars are the result. The declared result score decides each pair. Higher wins; exact equality draws. The rating chart is the replay of that result. It is not another score.

Standings

Result

Ordered by quality score. The highest score won. Equal scores are a draw. This is not a board rank. A positive rating change is not the same fact as winning the battle.

Tied highest score 5. Every system at that score is marked the same way.

  1. Claude Sonnet 5Anthropic via CursorExcellent against the assigned requirements, with no established material shortcoming in the available evidence5of 5
  2. Composer 2.5CursorExcellent against the assigned requirements, with no established material shortcoming in the available evidence5of 5
  3. GLM 5.2Z.ai via CursorExcellent against the assigned requirements, with no established material shortcoming in the available evidence5of 5
  4. GPT-5.6 SolOpenAIExcellent against the assigned requirements, with no established material shortcoming in the available evidence5of 5
  5. Grok 4.5xAIExcellent against the assigned requirements, with no established material shortcoming in the available evidence5of 5
  6. DeepSeek FlashDeepSeekStrong, well-supported work with limited gaps4of 5
  7. Kimi K3Moonshot via NVIDIAStrong, well-supported work with limited gaps4of 5
  8. Nemotron 3 Super 120B A12BNVIDIAStrong, well-supported work with limited gaps4of 5
  9. Grok 4.6xAIStrong, well-supported work with limited gaps4of 5
  10. Nemotron 3.5 Lightning 30B A3BNVIDIAUseful in part, but material shortcomings remain3of 5
Access route is not published

No separate access route is published for Composer 2.5, DeepSeek Flash, Nemotron 3 Super 120B A12B, Nemotron 3.5 Lightning 30B A3B, GPT-5.6 Sol, Grok 4.5, Grok 4.6. A missing route is not direct API access.

Selected contestant

Composer 2.5

Cursor

Model page

This battle: 5–0–4

5of 5. Excellent against the assigned requirements, with no established material shortcoming in the available evidence

Why this score

Explanation unavailable. The recorded result is unchanged.

Execution metrics

Measured statistics are not published for this contestant.

About this battle

Category
Quality
Difficulty
Not recorded
Published
Participants
10 participants

Rating impact (replayed)

Overall

1198.501204.69

+6.19

Quality

1199.261205.53

+6.27

A positive rating change is not the same fact as winning the battle.

Measured statistics are not published for this battle.

Pairwise comparisons

Higher quality wins. Equal quality draws. Filtering does not change the recorded outcomes.

4 draws

Replayed rating changes

Signed change first. Before and after are the replay arithmetic, not rank history. Shown movement is rounded to two decimals from the unrounded replay. It is not a saved historical publication.
SystemQuality score (1–5)Overall changeQuality change
Claude Sonnet 5 via Cursor5+6.131200.00 → +6.13 → 1206.13+6.241200.00 → +6.24 → 1206.24
Composer 2.55+6.191198.50 → +6.19 → 1204.69+6.271199.26 → +6.27 → 1205.53
GLM 5.2 via Cursor5+6.131200.00 → +6.13 → 1206.13+6.241200.00 → +6.24 → 1206.24
DeepSeek Flash4-3.981150.48 → -3.98 → 1146.51-4.441165.66 → -4.44 → 1161.22
Kimi K3 via NVIDIA4-5.871200.00 → -5.87 → 1194.13-5.761200.00 → -5.76 → 1194.24
Nemotron 3 Super 120B A12B4-5.871200.00 → -5.87 → 1194.13-5.761200.00 → -5.76 → 1194.24
Nemotron 3.5 Lightning 30B A3B3-12.021186.52 → -12.02 → 1174.51-12.361198.25 → -12.36 → 1185.89
GPT-5.6 Sol5+7.061175.77 → +7.06 → 1182.83+7.201175.10 → +7.20 → 1182.30
Grok 4.55+6.131200.00 → +6.13 → 1206.13+6.241200.00 → +6.24 → 1206.24
Grok 4.64-3.931149.38 → -3.93 → 1145.45-3.881150.69 → -3.88 → 1146.82

The signed change is the replayed update for this battle. Before and after are secondary. Overall and work-type changes stay separate. Shown movement is rounded to two decimals from the unrounded replay. It is not a saved historical publication.

Claude Sonnet 5

Anthropic via Cursor

Quality score (1–5)
5
Overall change
+6.131200.00 → +6.13 → 1206.13
Quality change
+6.241200.00 → +6.24 → 1206.24
Composer 2.5

Cursor

Quality score (1–5)
5
Overall change
+6.191198.50 → +6.19 → 1204.69
Quality change
+6.271199.26 → +6.27 → 1205.53
GLM 5.2

Z.ai via Cursor

Quality score (1–5)
5
Overall change
+6.131200.00 → +6.13 → 1206.13
Quality change
+6.241200.00 → +6.24 → 1206.24
DeepSeek Flash

DeepSeek

Quality score (1–5)
4
Overall change
-3.981150.48 → -3.98 → 1146.51
Quality change
-4.441165.66 → -4.44 → 1161.22
Kimi K3

Moonshot via NVIDIA

Quality score (1–5)
4
Overall change
-5.871200.00 → -5.87 → 1194.13
Quality change
-5.761200.00 → -5.76 → 1194.24
Nemotron 3 Super 120B A12B

NVIDIA

Quality score (1–5)
4
Overall change
-5.871200.00 → -5.87 → 1194.13
Quality change
-5.761200.00 → -5.76 → 1194.24
Nemotron 3.5 Lightning 30B A3B

NVIDIA

Quality score (1–5)
3
Overall change
-12.021186.52 → -12.02 → 1174.51
Quality change
-12.361198.25 → -12.36 → 1185.89
GPT-5.6 Sol

OpenAI

Quality score (1–5)
5
Overall change
+7.061175.77 → +7.06 → 1182.83
Quality change
+7.201175.10 → +7.20 → 1182.30
Grok 4.5

xAI

Quality score (1–5)
5
Overall change
+6.131200.00 → +6.13 → 1206.13
Quality change
+6.241200.00 → +6.24 → 1206.24
Grok 4.6

xAI

Quality score (1–5)
4
Overall change
-3.931149.38 → -3.93 → 1145.45
Quality change
-3.881150.69 → -3.88 → 1146.82

· Underlying task and output evidence is not published with this score record.

Shown movement is rounded to two decimals from the unrounded replay for this battle: before, change, after. It is not a saved point-in-time publication or seven-day trend. The board rounds the current rating. This page does not publish prompts or code.

More Quality battles