Benchmarks as Random Variables—Modeling Overdispersion in LLM Evaluation
Michal Buran ⋅ Václav Čadek
Abstract
Benchmark performance of large language models (LLMs) is typically reported as a simple average across tasks, collapsing task heterogeneity and run-to-run variability into a single number; in practice, repeated evaluations often exhibit more variability than binomial sampling predicts, especially in agentic settings. We model benchmark outcomes with a hierarchical Bayesian Rasch framework that separates agent ability from task difficulty, extended to a Beta-Binomial with a single concentration parameter $\kappa$ that captures overdispersion. Both on AIME 2025 and Toolathlon, the Beta-Binomial extension markedly improves fit consistent with high overdispersion. Although agent rankings remain stable ($\rho \ge 0.995$), standard approaches substantially underestimate uncertainty, with consequences for pairwise agent comparison and active evaluation.
Chat is not available.
Successful Page Load