One prompt, two hidden models

Every matchup presents two published game artifacts generated from the same versioned prompt. Candidate position is randomized, delivery links are unique to the matchup, and model identities remain hidden until a valid decision is recorded. An evaluator is not shown an artifact after its identity has been revealed to them.

What a vote means

Evaluators can choose candidate A, candidate B, or a tie after reviewing both games. Each evaluator can rate a canonical prompt and model pair once per season. A failed-game report removes that ballot from rating calculation and creates an operator-review signal; anonymous reports never remove content automatically.

Elo, without false precision

Elo estimates relative performance from head-to-head outcomes. New models are marked provisional until they have enough comparisons. Ratings are rounded for display, while deterministic ordering keeps equal ratings stable.

Auditability

Each accepted ballot references the exact prompt and output versions shown. It is accepted at most once, and rating changes are committed with the ballot so the leaderboard can be rebuilt from durable events.

Limits

This benchmark reflects the prompts, published artifacts, and people who participate. It does not measure every aspect of model quality. Generated games may also have accessibility limitations outside the platform shell.