Methodology
How the benchmark works
A practical explanation of blind comparisons, accepted ballots, and Elo ratings.
One prompt, two hidden models
Every matchup presents two published game artifacts generated from the same versioned prompt. Candidate position is randomized, delivery links are unique to the matchup, and model identities remain hidden until a valid decision is recorded. An evaluator is not shown an artifact after its identity has been revealed to them.
What a vote means
Evaluators can choose candidate A, candidate B, or a tie after reviewing both games. Each evaluator can rate a canonical prompt and model pair once per season. A failed-game report removes that ballot from rating calculation and creates an operator-review signal; anonymous reports never remove content automatically.
Elo, without false precision
Elo estimates relative performance from head-to-head outcomes. New models are marked provisional until they have enough comparisons. Ratings are rounded for display, while deterministic ordering keeps equal ratings stable.
Auditability
Each accepted ballot references the exact prompt and output versions shown. It is accepted at most once, and rating changes are committed with the ballot so the leaderboard can be rebuilt from durable events.
Limits
This benchmark reflects the prompts, published artifacts, and people who participate. It does not measure every aspect of model quality. Generated games may also have accessibility limitations outside the platform shell.