Feature Learning in High-Dimensions under Structured Covariance: Scaling Laws in Quadratic Networks
Abstract
Recent theoretical work has shown that nonlinear solvable models exhibit scaling laws in the feature-learning regime. However, these results largely rely on the assumption of isotropic inputs. Understanding how these laws extend to anisotropic data remains a central open problem. In this work, we address this question by analyzing the learning dynamics of two-layer neural networks with quadratic (second Hermite polynomial) activations under anisotropic Gaussian inputs. We provide a sharp characterization of stochastic gradient descent (SGD) across both continuous-time dynamics and finite-sample (online) discretizations, explicitly quantifying how the covariance spectrum influences the scaling exponent and sample complexity. Furthermore, we provide evidence that normalization techniques strictly improve sample efficiency in anisotropic settings. Experiments on two-layer networks with general activations support our theoretical predictions, suggesting that these insights extend well beyond the quadratic model.