A Mechanistic Audit of LLM Policy Decision-Making under Concept Injection
Abstract
Activation steering has been shown to alter large language model (LLM) behavior, yet little attention has been paid to structured policy decision-making contexts and the extent to which susceptibility varies between relevant and irrelevant injected concepts. We present a mechanistic audit of LLM decision-making in policy contexts, probing where decisions form and how susceptible they are to activation steering. We inject steering vectors into the model's hidden-state activations at selected transformer layers. The steering vectors encode either ethical relevance to policymaking or irrelevant themes to test LLM robustness to activation steering when prompted with synthetic policymaking scenarios. We find that policy decisions emerge predominantly in later transformer layers of an LLM, and irrelevant concepts flip decisions at rates matching or exceeding relevant ones. This indicates that LLM policy judgments are no more robust to semantically arbitrary manipulation than to grounded, semantically relevant steering. Furthermore, flip rates are moderated more strongly by policy framing than by concept category, with both categories producing comparable susceptibility across injection strengths. These findings inform our understanding of LLM susceptibility to targeted manipulation in high-stakes scenarios.