Belief Without Justification: Sycophancy as a Single-Layer Truth–Compliance Tension in LLMs
Valentin NOËL
Abstract
A sycophantic language model can state a fact correctly when asked neutrally, then disavow it the moment a user voices the opposite belief. Its behaviour tracks the user's belief rather than its own representation of the world, and the gap is invisible to output-level evaluation. We argue this is best understood as a tension instruction-tuning imposes on the same weights, truthfulness against compliance, and that the tension is mechanistically much more local than expected. In ten of eleven open-source instruction-tuned models (1B--14B parameters, six architectural families), sycophancy is mediated by the dominant singular subspace of a \emph{single} MLP weight matrix. Compressing that direction monotonically reduces sycophancy ($-31$ to $-77pp$); amplifying it induces sycophancy on otherwise neutral inputs ($+18pp$), establishing causal mediation. The eleventh model (Mistral-7B-v0.3, the only all-SWA architecture in our set) shows the opposite signature, predicted in advance from the SVD profile of its weights: compliance is distributed across the spectrum and surgery fails. The locus of a model's deference to users is a measurable, addressable structural property of its weights; tacit ``beliefs'' can be partially read off, and selectively edited, without retraining.
Chat is not available.
Successful Page Load