Leaderboard Incentives: Model Rankings under Strategic Post-Training
Abstract
Influential benchmarks incentivize competing model builders to strategically allocate post-training resources towards improvements on the leaderboard, a phenomenon dubbed \emph{benchmaxxing} or \emph{training on the test task}. In this work, we initiate a principled study of the incentive structure that benchmarks induce. We model benchmarking as a Stackelberg game between a benchmark designer who chooses an evaluation protocol and multiple model providers who compete simultaneously in a subgame given by the designer’s choice. Each competitor has a model of unknown latent quality and can inflate its observed score by allocating resources to benchmark-specific improvements. First, we prove that current benchmarks induce games for which no Nash equilibrium between model developers exists. This result suggests one explanation for why current practice leads to misaligned incentives, prompting model providers to strategize in opaque ways. However, we prove that under mild conditions, a recently proposed evaluation protocol, called tune-before-test, induces a benchmark with a unique Nash equilibrium that ranks models by latent quality. This positive result demonstrates that benchmarks need not set bad incentives, even if current evaluations do.
Lay Summary
Leaderboards play a major role in how large language models are compared, promoted, and improved. Because leaderboard rankings are based on benchmark scores, developers may have incentives to optimize specifically for those scores, rather than for broader improvements that matter in real-world use. This can make rankings misleading: a higher score may reflect targeted preparation for the test, not necessarily a better model overall. We study this issue by modeling benchmark evaluation as a game between the leaderboard designer and model developers. Our analysis shows that standard evaluation can create unstable incentives, where developers keep chasing small score improvements to move up the leaderboard. We then study a different approach called tune-before-test: before ranking models, the benchmark designer gives every submitted model the same standardized opportunity to adapt to the benchmark. In our model, this reduces the benefit of extra benchmark-specific optimization and helps the ranking better reflect the models’ underlying capabilities. The broader message is that benchmark design is also incentive design: good evaluations should consider how developers will respond to them.