Resolution Diagnostics for Paired LLM Evaluation
Anany Kotawala
Abstract
LLM leaderboards report percentage-point gaps but rarely the statistical resolution of those gaps under paired prompt dependence and leaderboard-scale multiplicity. We treat shared-prompt LLM evaluation as a paired hypothesis-testing problem: inverting level-$\alpha$, power-$(1-\beta)$ tests yields a minimum detectable effect, a required paired sample size $N^\star$, and a resolution ratio $q=N/N^\star$. A sharp non-asymptotic bound shows that the common unpaired Cohen-$h$-plus-$(1-\rho)$ shortcut underestimates $N^\star$ by a factor of two in the close-comparison regime, a deficit reproduced by widely used power-analysis tools. On the Open LLM Leaderboard v1, 11/40 unique pairwise rankings are unresolved at $(\alpha,1-\beta)=(0.05,0.8)$; on MMLU-Pro top-10, 4 of 9 adjacent-rank pairs are unresolved at $N=12{,}032$, rising to 6 of 9 under real subject-level clustering. The verdicts also survive multiplicity correction and anytime-valid sequential testing. Code: https://anonymous.4open.science/r/htw-llm-power-ED92/.
Chat is not available.
Successful Page Load