Sycophancy Is Often a Single-Layer Phenomenon: Spectral Diagnosis and Training-Free Weight Surgery in LLMs
Valentin NOËL
Abstract
Instruction tuned language models often defer to user opinions even when those opinions are factually wrong, a behavior known as sycophancy. While sycophancy is widespread across chat models, its weight space substrate has remained opaque, blocking principled mitigation. In this work, we show that sycophancy is mediated by the dominant singular subspace of a single MLP weight matrix in ten of eleven open source models spanning 1B--14B parameters. In each affected model a single layer concentrates the behaviour: compressing its dominant spectral direction monotonically reduces the forced choice sycophancy rate, and amplifying it induces sycophancy on otherwise neutral inputs, establishing causal mediation without contrastive data; aggressive compression at this layer reveals nonlinear capability tradeoffs that we analyze and operationalize as a layer geometry diagnostic. The per-layer spectral SNR does double duty: it identifies the target layer from model weights alone, and bounds the safe operating dose ($|\alpha| \ll 1/SNR_\ell$), a predictive criterion validated across model families before any behavioral evaluation. Leveraging this insight, we propose a closed form weight space intervention that requires no training, no contrastive data, and no inference time machinery, and that on Gemma-4-E2B-it produces a strict Pareto improvement: less sycophancy and more reasoning simultaneously. The per layer spectral profile partitions models into three storage classes, localised, weakly localised, and distributed, predicting in advance whether single layer surgery will succeed. Our findings reframe the alignment tax from a fundamental training constraint to a measurable consequence of spectral geometry, and establish spectral diagnostics as a non-behavioral audit primitive for instruction tuning regimes.
Chat is not available.
Successful Page Load