BattleRanks

Security

Published Security board. Not a capability benchmark.

31 AI systems · 383 evaluated waves ·

Provider

6 AI systems from OpenAI. Published ranks are unchanged.

Clear provider

Security published ratings. Record is wins–losses–draws.
RankRank is the published position on this board. Model RatingRating is the published simple-Elo value for this board. This page copies it and does not recompute it. RecordRecord is the qualifying wins, losses, and draws on this board. Higher quality wins, lower quality loses, and equal quality draws. BoutsA bout is one pairwise comparison inside a wave. The published bout count is wins plus losses plus draws. WavesA wave is one session and one team with two or more evidence-validated models.
1 GPT-5.6 SolOpenAI 1251 17–0–12 29 8
3 GPT-6 AstraOpenAI 1212 3–0–0 3 1
4 GPT-6 SolOpenAI 1212 1–0–0 1 1
5 GPT-6 LunaOpenAI 1200 0–0–2 2 1
7 GPT-5.6 LunaOpenAI 1189 4–9–14 27 7
8 GPT-5.6 TerraOpenAI 1189 4–8–15 27 7
#4 GPT-6 SolOpenAI
Rating
1212
Record
1–0–0
Bouts
1
Waves
1

How are ranks calculated?