On How Muon Reshapes Skill Learning Dynamics
Abstract
Spectrum-aware optimizers, particularly Muon, exhibit more favorable empirical scaling than optimizers like SGD and Adam. Prior work explains this through more balanced learning in settings with imbalanced labels or input-output associations. In this work, we show that Muon enables more balanced skill acquisition in settings with task-level imbalance. Using a multi-task sparse parity-based setup, we show that with Muon, sub-tasks are learned more in parallel than with other optimizers. We introduce a Skill Acquisition Lag (SAL) metric to quantify parallel versus sequential learning. In a simple in-context linear regression setting, we show that Muon's convergence is independent of both task- and input-level imbalance, in contrast to (normalized) GD.