How Language Models Choose Sides: Internal Representations of Instruction Hierarchy
Enrique Balp-Straffon ⋅ Chih-Hao Hsu ⋅ Rushiraj Gadhvi ⋅ Sunishchal Dev ⋅ Callum McDougall ⋅ Anusha Mujumdar
Abstract
We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints with deterministic verifiers and evaluate eight models under matched baseline, conflict, and same-channel control conditions.Behaviourally, the models split into three regimes by System Authority Delta: hierarchy-respecting models use the system channel as an authority signal, anti-hierarchy models follow the system less often than their same-channel baseline predicts, and no-effect models show little channel sensitivity. Llama-3.1-8B is the strongest anti-hierarchy case in our suite, following the system in only $0.10$ of conflict trials.We use this behavioural failure case to ask whether user-preferring arbitration reflects the absence of an internal conflict-resolution signal. It does not: on Llama-3.1-8B, the conflict outcome is linearly decodable from residual-stream activations at $0.97$ balanced accuracy, $17$ percentage points above a metadata-only baseline, with analogous signals on Qwen2.5-7B and gpt-oss-20b.Steering with a layer-$12$ mean of four per-conflict logistic-regression directions raises genuine system compliance from $0.132$ to $0.530$, while directions selected mainly for pooled separability steer poorly. User-preferring conflict resolution can therefore coexist with a readable internal arbitration signal, and successful intervention depends on the geometry of the readout rather than probe accuracy alone. Code is available at https://anonymous.4open.science/r/system-user-circuits.
Chat is not available.
Successful Page Load