Biregular Sparse Initialization Shifts the Rate and Shape of Compositional Escape in Sequential Arithmetic Curricula
Clément Castellon ⋅ Arindam Biswas
Abstract
Sequential curriculum learning on modular arithmetic produces a characteristic interference ceiling: a 2-layer transformer trained to generalize addition and then exposed to multiplication typically plateaus near 43\% validation accuracy and remains there. We study this ceiling under biregular sparse initialization using Ramanujan graphs and matched unstructured sparse controls, testing three preregistered hypotheses. Our main finding is that biregular sparse structure shifts the rate of compositional escape beyond matched Bernoulli sparse controls: at sparsity 0.977, Ramanujan initialization escapes in 11/15 runs ($\geq 0.95$ validation accuracy) versus 4/15 for unstructured sparse at the same density (Fisher exact test: $p = 0.013$ one-sided, $p = 0.027$ two-sided), while unstructured escape rate is flat across all sparsity levels. Biregular initialization also changes the \emph{shape} of the outcome distribution: unstructured networks produce ambiguous partial escapes (67\% of runs in the 0.50--0.95 range at sparsity 0.977), while Ramanujan networks shift mass away from ambiguous partial escapes and toward clean failure or full escape. A spectral diagnostic shows that the Ramanujan margin is nearly zero at high sparsity ($\lambda_2 \approx$ bound at 0.977), ruling out the margin as the operative mechanism; we hypothesize that the advantage instead comes from degree uniformity, which guarantees nonzero row and column degree by construction. An embedding ablation shows that embedding dynamics contribute substantially to phase-2 escape across all conditions; freezing embeddings reduces escape rates for dense, Ramanujan, and unstructured models alike. Biregular weight structure does not measurably compensate for removing embedding dynamics, leaving the mechanism of the structural advantage in the natural curriculum open. We document a threshold-collapse artifact that inflates magnitude-based circuit overlap on sparse networks with externally stored mask buffers.
Chat is not available.
Successful Page Load