Sham M. Kakade (Harvard), How far does Taylor's Theorem take us in deep learning?
Abstract
Continuing from Wednesday, we take the local quadratic approximation seriously and ask how well it holds during LLM pretraining. Taylor expanding real training runs around their iterates, we find — somewhat surprisingly — that linearized models track the true loss for a meaningful window and that this continues to hold deep into training. Using Lanczos quadrature, we then compute the full spectral density of 150M-parameter models, finding power-law spectra whose curvature is dominated by identifiable parts of the network. Together, these suggest Taylor's theorem is a more powerful lens on deep learning than its simplicity implies.
Video
Chat is not available.
Successful Page Load