Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
Abstract
Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimensions. It was also shown that this effect persists in the limit of highly overparametrized networks as well as the Neural network Gaussian process limit. We here develop a principled understanding of the typical behavior of generalization in Neural Network Gaussian process regression dynamics. We derive a dynamical mean-field theory that captures the typical case learning dynamics: This allows us to unify multiple existing regimes of learning studied in the current literature, namely Bayesian inference on Gaussian processes, gradient flow with or without weight-decay, and stochastic Langevin training dynamics. Employing tools from statistical physics, the unified framework we derive in either of these cases yields an effective description of the high-dimensional microscopic behavior of networks dynamics in terms of lower dimensional order parameters. We show that collective training dynamics may be separated into the dynamics of N independent eigenmodes, whose evolution equations are only coupled through collective response functions and a common statistics of an effective, independent noise. Our approach allows us to quantitatively explain the dynamics of the generalization error by linking spectral and dynamical properties of learning on data with power law spectra, including phenomena such as neural scaling laws and the effect of early stopping.
Lay Summary
When we train large artificial neural networks, their ability to make accurate predictions on new, unseen examples improves in a strikingly regular, predictable way as we add more data or make the networks bigger. These patterns, called "scaling laws," seem to hold no matter how the network is built — but why they emerge has remained poorly understood. We developed a mathematical theory, borrowing tools from physics, that describes how a typical neural network learns over the course of training. Our theory unifies several different ways of training networks that researchers had previously studied separately, showing they are all special cases of one common picture. This lets us explain, from first principles, why scaling laws appear and how training choices — such as stopping early to avoid overfitting — affect performance. By turning these widely observed patterns into something we can understand and predict, our work offers a foundation for training models more efficiently.