Blake Bordelon (UT Austin), Training Dynamics in Large Networks: From Super-Wide to the Scaling Law and Transfer Regime
Abstract
In this talk, we discuss stable scaling of deep learning models by identifying stable, feature-learning infinite width and depth limits of neural networks. The asymptotic description of randomly initialized networks in this regime will take the form of a dynamical mean field theory (DMFT). We will discuss how adoption of scaling strategies that admit such limits yields better hyperparameter transfer, where optimal hyperparameters in small models remain optimal in large models. We will provide examples of these results for multi-layer perceptrons, convolutional networks, self-attention blocks, and mixture-of-experts transformers. These exact limits assume parameters are much larger than the training data, and thus fail to capture the behavior of models in the scaling law regime where models of different sizes achieve different performance. To address this, we will introduce simplified, analytically tractable models which enable analysis of training dynamics far from the infinite limit. We show that early time deviations in model performance are universal, while late time deviations are architecture and data dependent. We use one of these toy models to analyze hyperparameter transfer across training horizons T through an optimal control paradigm, showing that the optimal strategy is highly dependent on feature structure and SGD noise statistics.