Benchmark trust
How the rankings work
Effective August 10, 2026
One task, four anonymous candidates
Each arena resolves the user’s brief into a writing category and sends the same serialized task, context, and generation policy to four independently selected model routes. Model names remain hidden until the battle is complete, and candidate positions are secret-key permuted.
Four decisions, one clear bracket
Two semifinals send their winners to the championship and their other candidates to a third-place match. Third and fourth place are settled first; the championship is always the final decision and determines first and second. People may choose A, B, or a tie and may separately flag that both candidates need work. Ties contribute half a win to each model; a precommitted, pair-order-independent tiebreak advances only the bracket path.
Eligible evidence only
Published snapshots use completed arenas from authenticated, email-verified people. Extremely fast decisions, quarantined activity, incomplete model fields, and excess weekly contributions are excluded. Per-person and per-pair caps prevent a small number of prolific users from dominating a model’s score.
Category-specific Elo
WriteArena fits a regularized Bradley–Terry preference model separately for creating and improving writing, both overall and by category. Scores are centered at 1000 Elo, and every model in the active catalog remains visible while the model learns from new comparisons. Low-volume models stay pulled toward the shared center rather than disappearing behind an arbitrary publication floor.
Win rates and error bars
Overall win rate is preference share across eligible comparisons, with ties contributing half a win. Direct model-versus-model results use only comparisons where that exact pair met. Their error bars are approximate 95% Wilson intervals; leaderboard Elo intervals come from 500 user-clustered bootstrap refits, which preserve the fact that one person may cast multiple votes.
Quiet interface, complete record
The default leaderboard keeps the live view readable: rank, Elo, win rate, and each model’s win–loss–tie record. Full counts, confidence intervals, eligibility fields, and reproducible snapshot metadata remain available in the machine-readable benchmark data.
Reproducible snapshots
Every benchmark era freezes the model catalog, provider route, prompt harness, category rules, generation policy, comparison inputs, solver version, and eligibility policy. Published runs store dataset and result hashes so a chart can be traced back to the exact evidence that produced it.
What the benchmark does not claim
WriteArena measures preference for writing produced in its own task mix and interface. It is not a general intelligence score, a guarantee for every prompt, or a substitute for reading the outputs. Rankings are most useful when the selected mode, category, evidence level, and snapshot date match the decision you are making.
