Spectral Equalization Minimizes Total Training Energy: A Control-Theoretic Account of Muon's Advantage
Abstract
We propose the total training energy, the integral squared error of the back-propagated signal, as a trajectory-level diagnostic for deep learning optimizers, viewing every first-order method as a discrete-time feedback controller closing the loop around the loss landscape. On quadratic losses with a Kronecker-factored Gauss–Newton Hessian, the total training energy equals the squared H2 norm of the closed loop and decomposes exactly into per-mode energy contributions. Jensen’s inequality implies that uniform per-mode rates strictly minimize the total training energy among controllers with a common weighted-mean contraction rate, which Muon’s polar step implements to the extent its update directions align with the Kronecker eigenbasis. On MNIST with one- , two-, and three-layer MLPs trained with SGD, AdamW, and Muon, Muon attains the smallest Gini coefficient of the per-mode energy distribution on every monitored hidden matrix and the smallest cumulative energy on two of three, with the advantage concentrated in the low-curvature tail of the spectrum. The framework recasts “orthogonalizing momentum helps” as a measurable, mechanism-level claim about the geometry of the Hessian spectrum.