Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization
Jiajie Zhao ⋅ Jianxing Wang ⋅ Junjie Yang ⋅ Zhiwei Bai ⋅ Yaoyu Zhang
Abstract
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified $\mathcal{l}_1$ norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified $\mathcal{l}_1$ norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.
Lay Summary
(1) Imagine giving a computer a riddle with millions of correct answers. Miraculously, the AI chooses a simple answer. Scientists call this built-in preference "implicit bias." But why does AI have such great intuition? What unseen forces guide it to the best solution? (2) To crack this puzzle, we analyzed a simplified AI model known as a diagonal linear network. By starting the model from a nearly blank slate (infinitesimal initialization), we developed an algorithm to map out its exact learning trajectory. Besides, we identified an intrinsic geometric structure—which we call Structural Invariant Manifold—that shapes how the model's parameters evolve. (3) We proved that the model’s learning process naturally forces it to seek out solutions with a minimal modified $\mathcal{l}_1$ norm. Ultimately, this discovery deepens our fundamental understanding of implicit bias in artificial intelligence.
Successful Page Load