A Training-Dynamics View of Catastrophic Overfitting: Understanding and Prevention
Abstract
We study the underlying mechanism of catastrophic overfitting, a phenomenon in which models overfit to weak adversarial examples and lose true robustness, from the perspective of training dynamics. Despite numerous studies investigating catastrophic overfitting, the fundamental cause of sudden robustness collapse remains poorly understood. In this study, we systematically analyze the dynamics of adversarial training and reveal that the rapid amplification of the \emph{mixed Hessian} causes catastrophic overfitting. Based on this insight, we propose a novel KL divergence-based regularizer that stabilizes training dynamics and effectively prevents catastrophic overfitting. Remarkably, our method consistently matches or even surpasses the robustness of multi-step adversarial training, despite using single-step adversarial training. Furthermore, when combined with multi-step adversarial training, our regularizer yields additional robustness improvements, indicating that mixed Hessian stabilization is a general principle applicable beyond the single-step regime.