Practical Muon Accelerates Projected Feature Learning in Scaling-Law Models
Abstract
We ask whether optimizer geometry can improve feature-learning scaling laws in the two-layer linear student-teacher model of Bordelon et al. (2025). Starting from their projected feature-learning dynamics, we replace the feature update with practical Newton-Schulz Muon while preserving the finite-width projection bottleneck. As a proxy theory, we derive a projected partial-polar population flow and show that it removes the small projected-gradient norm factor inherited by projected SGD on hard source-condition tasks. Empirically, practical NS5 Muon yields positive fixed-compute width-scaling exponents on hard tasks, whereas projected SGD has negative fixed-compute width exponents over the same sweep. Readout-Hessian diagnostics show a smaller 0.99-quantile gap decay exponent, consistent with stronger low-rank feature alignment. These results suggest that matrix orthogonalization can materially change finite-compute feature-learning scaling.