A loss curvature account of fine-tuning fragility
Abstract
Fine-tuning on narrow distributions often produces fragile changes that are easily reversed by further training, with implications for the durability of safety fine-tuning. Mixing pre-training data into fine-tuning is a known mitigation, but why varying the proportion of fine-tuning data (which we term concentration) modulates forgetting is poorly understood. During a reversion phase (subsequent training on pre-training data after fine-tuning), we decompose the per-step change in fine-tune loss into its first- and second-order Taylor terms. We then track how each varies with concentration. In experiments on LLMs (Pythia-70M), we find that the second-order (curvature) term grows in importance with concentration. This curvature is much larger along the reversion update direction than along a random direction at every concentration, with the directional curvature itself increasing with concentration. Curvature therefore contributes to forgetting even when fine-tune and pre-train gradients are not in conflict, and its relative share grows with concentration, providing empirical support for recent theoretical accounts of curvature-driven forgetting.