The Geometry of Forgetting: A Fisher-Information Framework for Alignment-Preserving Continual Adaptation
Soham Batra ⋅ Siddharth Karuturi ⋅ Kaustubh Bukkapatnam ⋅ Laksh Patel
Abstract
Continual adaptation of aligned foundation models risks progressive degradation of safety-critical behaviors---a phenomenon we call \textit{alignment forgetting}. We provide the first rigorous geometric characterization of this process via the Fisher information matrix of the alignment distribution. Concretely, we show that alignment degradation incurred by any parameter update $\delta$ equals exactly $\frac{1}{2}\delta^{\top}F_{A}\delta$ (under standard second-order approximation), and that alignment-safe adaptations lie in a computable Fisher ellipsoid $\mathcal{S}_{\epsilon}$. Building on this, we derive a closed-form KKT expression for the \textit{alignment tax}---the unavoidable task-performance cost of respecting alignment constraints---and a quadratic law governing how alignment erodes over sequential adaptation steps. Finally, we formulate optimal deliberate forgetting (machine unlearning) as a generalized eigenproblem in Fisher space, yielding a principled algorithm we call \textit{FAE} (Fisher Alignment-Ellipsoid). Experiments on LLaMA-3-8B across 30 tasks confirm our predictions ($R^2=0.991$ for predicted vs. actual degradation; diagonal Fisher approximation quality $R^2=0.887$ vs. full Fisher, see Appendix E), and FAE outperforms strong unlearning baselines on all three axes of forget quality, knowledge retention, and alignment preservation on the TOFU benchmark.
Video
Chat is not available.
Successful Page Load