SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR
Siddharth Aphale ⋅ Kelly Liu
Abstract
The standard recipe of picking the highest scoring SFT checkpoint for GRPO can pick the worst one when the policy has entered entropy collapse during SFT. We show that under binary rewards the within group advantage variance is exactly $p(1-p)(g-1)/g$, and characterise a threshold $p^{\ast}(g)$ below which the GRPO gradient vanishes. We test this on two SFT depth ladders, Qwen2.5-Coder-3B and DeepSeek-Coder-6.7B, across 1.0--5.8 epochs and 3 seeds. The Qwen ladder crosses $p^{\ast}(8)$ at deeper checkpoints and shows rank inversion: peak GRPO pass@10 falls monotonically from 0.806 to 0.481 (3 seed mean, $n=20$), with pre-RL pass@1 correlating inversely with GRPO outcome ($\rho=-0.75$) and pre-RL entropy correlating positively ($\rho=+0.69$). DeepSeek stays $\geq 4.2\times$ above $p^{\ast}(8)=0.083$ and shows rank compression with uniformly positive $\Delta$pass@10, providing a contrastive validation of the threshold prediction; $p^{\ast}(g)$ correctly classifies both experimental regimes before any RL compute is spent. A two stage diagnostic, a pre-RL entropy screen and a 150 step in training monitor, rejects high risk checkpoints before GRPO (saving 100\% of RL compute) or by step 150 of 400 (saving 62.5\%), and improves deep eval pass@10 by +0.090 ($n=128$) when used for initialisation. Neither KL regularisation nor label smoothing eliminates the rank inversion, ruling out hyperparameter and RL-stage artefacts as explanations.
Chat is not available.
Successful Page Load