Rank-One Spectral Normalization Accelerates Projected Feature Learning
Chryseis Liu ⋅ Manuel Paez
Abstract
We ask whether optimizer-induced matrix normalization can improve feature-learning scaling laws in the projected two-layer linear student-teacher model of Bordelon, Atanasov, and Pehlevan. In this model, after the fixed finite-width projection bottleneck, the population feature-gradient matrix is rank one. Projected SGD therefore scales its feature update by the singular value of this rank-one matrix, while polar normalization removes that singular-value factor. We show that zero-momentum practical NS$_5$ Muon reduces exactly to this polar direction up to a fixed scalar in the rank-one population setting. Empirically, practical NS$_5$ Muon yields positive fixed-compute width-scaling exponents on hard source-condition tasks where projected SGD has negative exponents. Mechanism controls inspired by recent critiques of Muon-specific geometry show that exact polar updates, Frobenius-normalized gradients, norm-matched random spectra, and Freon with $c=1/2$ reproduce Muon's projected-population scaling once this scalar is matched. Thus the evidence supports a broad rank-one spectral-normalization mechanism, with practical Muon as an efficient implementation, rather than a Muon-unique geometric effect.
Chat is not available.
Successful Page Load