A Quadratic Lens on Muon: Orthogonalization, Invariance, and Implicit Preconditioning
Abstract
Muon and related optimizers improve training by approximately orthogonalizing matrix-valued updates, but the geometry behind their empirical gains remains partially understood. We study an idealized polar version of Muon on matrix-quadratic objectives, using exact line search to isolate update direction from step-size effects. This exposes a gain–curvature tradeoff with three exactness regimes: curvature isotropy (GD), coordinate-aligned vertex structure (SignGD), and row-isotropic iterates (PolarGD), the last holding independently of Hessian conditioning. We show PolarGD is equivariant under orthogonal basis changes, unlike SignGD, explaining its robustness to rotations of ill-conditioned structures in Zipf-style regression experiments. Beyond one step, PolarGD admits an implicit iterate-dependent preconditioner that whitens iterate covariance in the full-row-rank regime, yielding an exact progress coefficient separating the fast row-isotropic regime from the pessimistic tiny-singular-value regime. Rank-deficient extensions replace global by active-subspace curvature, and finite Newton–Schulz orthogonalization is shown to act as a bounded spectral filter that damps exact polar updates near rank deficiency.