On the Geometric Structure of Token–Token Interactions in Deep Language Models
Abstract
We introduce an operational framework for measuring and controlling token--token interactions in language models through controlled perturbations of non-target tokens. Holding the target token and syntactic frame fixed, we vary contextual tokens systematically and measure how the target representation changes. This setting lets us identify interaction effects directly, rather than inferring them from attention weights, architecture-specific components, or raw hidden-state contrasts. Our key idea is to residualize layer updates with respect to a restricted tokenwise approximation class. We decompose each target-token update into a token-local component and a residual, then ask whether category-conditioned variation induced by controlled perturbations concentrates more strongly in the residual. Across semantic categories, targets, and syntactic templates, we find that it does: residual directions capture structured interactional variation more cleanly than raw update directions, while directions derived without tokenwise subtraction retain token-local contamination that degrades, and in some cases reverses, steering precision. We then test whether these residual directions are causally meaningful. Injecting category-specific residual directions produces monotonic, category-selective changes in model outputs in settings where the interaction signal is cleanly isolated. We validate this framework across Transformer (Pythia, LLaMA) and state space (Mamba) models from 160M to 70B parameters, showing statistically significant interaction structure under permutation testing, reliable causal steering through residual injection, and category-conditioned transport consistency across target tokens. Together, these results support residualized layer-update geometry as a practical, architecture-agnostic interface for analyzing and intervening on contextual computation.