A "feature ODE" describing the learning behavior of shallow MLPs on simple functions
Joseph Turnbull ⋅ Berkan Ottlik ⋅ James B Simon
Abstract
We study the gradient flow training dynamics of two-layer multi-layer perceptrons (MLPs) on anisotropic Gaussian data. Even in this simple setting, solving the dynamics in full is notoriously difficult, so we try something different. We propose a "feature ODE" that might be intuitively expected to describe the learning dynamics of a shallow MLP. We find that MLPs trained to learn simple target functions (low-order Hermite polynomials and staircase functions) actually roughly follow the loss trajectories from this feature ODE. This agreement is a surprise because of the size of the simplification and the fact that it does not come from any rigorous chain of approximations. We discuss this surprise and speculate on its reasons and uses.
Chat is not available.
Successful Page Load