BattleRanks

Product

Published Product board. Not a capability benchmark.

31 AI systems · 383 evaluated waves ·

Provider

5 AI systems from OpenAI. Published ranks are unchanged.

Clear provider

Product published ratings. Record is wins–losses–draws.
RankRank is the published position on this board. Model RatingRating is the published simple-Elo value for this board. This page copies it and does not recompute it. RecordRecord is the qualifying wins, losses, and draws on this board. Higher quality wins, lower quality loses, and equal quality draws. BoutsA bout is one pairwise comparison inside a wave. The published bout count is wins plus losses plus draws. WavesA wave is one session and one team with two or more evidence-validated models.
3 GPT-5.6 LunaOpenAI 1210 3–2–3 8 5
4 GPT-6 AstraOpenAI 1206 1–0–1 2 1
6 GPT-6 LunaOpenAI 1202 3–2–1 6 1
8 GPT-5.6 SolOpenAI 1200 0–0–1 1 1
10 GPT-5.6 TerraOpenAI 1197 9–9–11 29 17

How are ranks calculated?