Quantifying Ranking Uncertainty in LLM Benchmarks
Bitya Neuhof ⋅ Yuval Benjamini
Abstract
Pretrained models are typically evaluated on multi-task leaderboards to assess their effectiveness across diverse tasks. Although recent work has introduced interval-based model rankings, variation in model performance across tasks and the interpretation of ranking uncertainty are often overlooked. In this work, we analyze sources of uncertainty in the MMLU benchmark and show how the hypothesis tests that form the ranking intervals can be modified to study these different sources. We show that variability in rankings across subjects is substantial and must be considered when comparing LLMs or identifying the top-performing models.
Chat is not available.
Successful Page Load