The Self-Verification Cliff: Generation Outpaces Self-Selection in Frontier LLMs, and It Widens With Capability
Arihant Jain
Abstract
Self-evolving scientific agents must do more than generate candidate solutions: they must select or verify which of their own outputs to trust, because every self-play, self-rewarding, and reinforcement-learning loop that omits an external checker implicitly assumes the model can grade itself. We test that assumption on competition mathematics, where final answers are integers and grading is therefore an exact, noise-free oracle. Across two frontier model families on 86 AIME problems, we find a large and systematic self-verification cliff: a model contains a correct answer among $k{=}8$ samples far more often than it can pick that answer using its own judgment. GPT-5.4-mini reaches a correct solution on $54.7\%$ of problems but self-selects correctly on only $41.9\%$, recovering just $36\%$ of its sampling headroom; Gemini-3.5-flash recovers $43\%$. Critically, the cliff widens with generation ability rather than closing: the stronger generator has the larger absolute gap. We further show that an external verifier helps only when it is at least as capable as the generator---a stronger cross-family verifier closes part of the cliff ($+3.5$ pts), while a weaker one makes selection worse than self-selection ($-2.3$ pts)---and that majority voting fails to rescue selection on these diverse-answer problems. The cliff is a structural ceiling on same-model self-improvement and a concrete argument for external (e.g.formal) verification in self-evolving math agents.
Successful Page Load