Invited Talk #3: Rotation Invariance, Saddle Points, and Linear Structure in LLMs: Implications for Quantization
Abstract
Large language models exhibit a surprising robustness to a range of transformations in weight space that, in principle, should significantly alter their behavior. These include drastic changes in numerical precision, as well as seemingly unrelated operations such as rotations and linear interpolation between pretrained models. Understanding why these transformations often preserve performance remains an open question.
In this talk, I propose a unifying perspective: these phenomena reflect an underlying structure in the weight space of LLMs, which we refer to as a weight-space transformation invariance. We demonstrate that representation geometry induces a strong form of rotation invariance, whereby rotationally equivalent reparameterizations of a full-precision neural network can exhibit markedly different robustness to quantization. We also find that the optimization landscape is dominated by saddle-point structures that shape the behavior of training under perturbations. In addition, we observe that pretrained models exhibit unexpectedly strong linearity in weight space, enabling effective interpolation and mixing between models.
Together, these findings suggest that LLM has a deep geometric and algebraic structure in neural weight spaces. Characterizing this structure provides a more general foundation for understanding and designing transformations of large language models.