The learning dynamics of geometric structure in transformers
Marmik Chaudhari
Abstract
Language models trained on various tasks exhibit a rich $\textit{geometric structure}$ in the representation space, such as smooth manifolds for counting characters in a line. But how these geometric structures $\textit{emerge}$ during training, when they become $\textit{useful}$ for computation, and how their formation is reflected in the model $\textit{weights}$ remain poorly understood. In this work, we study these learning dynamics in a small transformer trained to predict line breaks in synthetic text at a fixed line width. Specifically, we find a $\textit{phase change}$ that is primarily driven by the emergence of a 1D ordinal manifold for character counting, coinciding with the onset of accurate line-break prediction. This geometric structure exerts a progressive causal influence on the task. Early in training, ablating the character count subspace has no effect, but as the manifold forms and converges towards its final geometry, the same ablation increasingly degrades loss and suppresses line-break token prediction. At the weights level, we show that this phase transition exhibits the signature of $\textit{saddle escape}$: low-rank updates in the attention weights and a transient negative-curvature direction in the loss landscape. Finally, we sweep the weight initialization scale $\alpha$, which monotonically slows the phase transition as $\alpha$ decreases. More broadly, our work highlights the geometric structure as a lens for understanding how transformers acquire task-relevant representations during training.
Chat is not available.
Successful Page Load