A loss curvature account of fine-tuning fragility
Abstract
Fine-tuning on narrow distributions often produces fragile changes that are easily reversed by further training, with implications for the durability of safety fine-tuning. Mixing pre-training data into fine-tuning is a known mitigation, but why varying the proportion of fine-tuning data (which we term concentration) modulates forgetting is poorly understood. During a reversion phase (subsequent training on pre-training data after fine-tuning), we decompose the per-step change in fine-tune loss into its first- and second-order Taylor terms. We then track how each varies with concentration. In experiments on LLMs (Pythia-70M), we find that the second-order (curvature) term grows in importance with concentration, and that this sharpness lies specifically along the reversion update direction, growing monotonically with concentration. Curvature can therefore erase fine-tuned behaviour even when fine-tune and pre-train gradients are not in conflict, providing empirical support for recent theoretical accounts of curvature-driven forgetting.