Rank Allocation in Low-Rank Optimizers
Ansh Tiwari
Abstract
Every large model trained today moves a per-step rank budget across its layers through a low-rank spectral optimizer, Muon, Dion, PowerSGD, or GaLore, and in every case the split of that budget across attention and feed-forward roles is a one-line constant, set once and never inspected. We formalize this family as \emph{rank-profile spectral optimizers}, separate its two natural capture geometries (Ky-Fan partial sums for orthogonalized descent, Frobenius partial sums for low-rank compression), and prove a matched-budget lower bound that pins down, in measurable per-layer spectral quantities, how much captured signal a uniform allocation surrenders to a role-conditioned profile; the central structural inequality is machine-checked in Lean~4. A matched-budget probe in a $166$M-parameter LLaMA decoder at $d_{\mathrm{head}}=128$ ($3$ paired seeds) resolves the rank-allocation surface to within paired-seed noise: a rank-inverted profile loses to the uniform baseline by $+0.033$ nats (95\% paired-bootstrap CI $[+0.027,\,+0.041]$, $3/3$ sign), and the published $(1,4,1)$ structural constants land between inverted and uniform at $+0.015\ [+0.009,\,+0.026]$. Role assignment is load-bearing inside a fixed rank budget, and the operating point implied by the spectrum is not the one currently in use.
Chat is not available.
Successful Page Load