SIB: Reparameterization of LLMs for Better Learning-Forgetting under SFT
Abstract
Supervised finetuning (SFT) of pretrained language models trades off the acquisition of new domain capabilities against retention of prior knowledge. Recently, post-training quantization (PTQ) and catastrophic forgetting from finetuning are increasingly seen as a loss geometry problem, where flatness leads to lower degradation. In this work, we adopt a unified view of post-training perturbations. In particular, inspired by PTQ we propose \textbf{Scale Invariant Balancing (SIB)} a functionally equivalent reparameterization within the weight-space symmetries that flattens the loss landscape. We extensively characterize the learning-forgetting trade-off for SFT and SIB. Across models and methods, two regimes universally develop. Either baseline SFT performance appears as a gradual trade-off between learning and forgetting, in which case SIB can be applied to approximate Pareto optimality, or, baseline SFT is already not forgetting, in which case SIB does not substantially intervene.