Emergent Structured Representations Support In-Context Inference in Large Language Models
Abstract
Whether large language models (LLMs) perform structured inference or merely exploit surface-level statistical associations remains a central debate. Here, rather than dissecting individual attention heads or neurons, we isolate a shared latent subspace in the residual stream and test its functional role in inference. Our results show that this subspace emerges in intermediate layers and exhibits increasing cross-context alignment with more demonstrations. Causal interventions reveal that it is not a mere epiphenomenon but a functional mediator of inference: restoring it recovers model performance under corruption, ablating it significantly degrades performance, and transferring its relational structure enables targeted steering of model predictions across contexts. We further identify a layer-wise progression where attention heads in early-to-middle layers integrate contextual cues to construct the subspace, and later layers leverage it to generate predictions. Together, these findings suggest that LLMs dynamically assemble a structured and transferable latent substrate for inference, offering insights into the computational foundations of flexible adaptation and a promising basis for steering model behavior.