Do LLMs “Feel”? Emotion Circuits Discovery and Control
Abstract
Lay Summary
Can we locate emotion inside a language model the way neuroscientists map functional regions in the brain? This is the question that motivated our work. Prior research has shown that language models encode emotion-related signals in their internal representations, but nobody has known which specific components produce these signals or how they cooperate to give rise to emotional expression. We constructed a controlled dataset in which all input texts are deliberately free of emotion words, ensuring that when the model is guided to express different emotions, any differences in internal activations can only be attributed to the emotion mechanism itself rather than to the input content. Building on this, we extracted emotion representation directions that remain stable across different contexts, then used analytical decomposition and causal intervention to precisely locate which neurons and attention heads drive each emotion. Finally, we integrated these locally identified components from across different layers into coherent global emotion circuits, and directly modulated these circuits to control the model's emotional expression. On LLaMA-3.2-3B-Instruct, this achieves over 99% accuracy in inducing all six basic emotions without any emotion-related instructions in the prompt. This work demonstrates that emotional expression in language models is supported by internal mechanisms that can be located, interpreted, and intervened upon. Identifying these mechanisms brings us a step closer to genuinely understanding the emotional behavior of AI systems, and provides a methodological foundation for transparent auditing and targeted safety interventions on emotional expression in the future.