Reset-and-Discard (ReD) Improves Coverage at every Budget under Inference Power-Law Scaling
Sagi Meir ⋅ Tommer D Keidar ⋅ Noam Levi ⋅ Shlomi Reuveni ⋅ Barak Hirshberg
Abstract
The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@$k$, the probability of answering a question correctly at least once in $k$ trials. At a fixed budget across a workload of many tasks, a more suitable metric is coverage@cost: the expected number of unique questions answered as a function of total attempts. We connect these metrics via renewal theory and show that the empirically-observed power-law scaling of pass@$k$ (with exponent $0<\alpha<1$) leads to sublinear (diminishing-returns) growth of coverage@cost under standard solve-to-completion allocation. We propose Reset-and-Discard (ReD), a cross-problem allocation policy that provably restores linear coverage growth and maximizes coverage@cost at every budget, even under imperfect verifiers. ReD also provides a statistically efficient method to estimate inference power-law exponents when large $k$ pass@$k$ measurements are expensive. Experiments across three LLMs and three benchmarks show large reductions in required attempts, tokens, and USD cost.
Chat is not available.
Successful Page Load