The Channel Geometry of Refusal: Mechanistic Diagnosis of Alignment Collapse Under KV Quantization
Bruce C Xu ⋅ Adarsh Kumarappan ⋅ Mu Zhou
Abstract
Refusal in instruction-tuned LLMs is mediated by a small set of activation directions concentrated in the earliest output tokens. We use KV cache quantization as a controlled probe of where those directions live and how robust they are to per-channel rounding noise. Across eleven instruction-tuned models (3.8B-72B), low-bit KV quantization triggers sharp phase transitions in refusal invisible to perplexity monitoring: Mistral-7B loses 15.2\% of its refusals at only $1.03\times$ perplexity. The collapse is explained by a single channel-geometry property (whether the channels carrying refusal overlap the activation outliers a quantizer must accommodate), captured in a closed-form bound on the MSE gap between per-tensor and per-channel quantization. The empirical realization, **Per-Channel Reduction** (PCR), is a 20-prompt diagnostic that sorts models into three failure modes (*outlier-crushes-safety*, *outlier-as-safety*, *multi-layer dilution*). Read together, these modes reveal that the apparent disagreement between single-direction and multi-orthogonal accounts of refusal is not a contradiction but two endpoints of a concentrated-to-distributed spectrum whose position is determined not by architecture but by the post-training recipe. The resulting protocol recovers up to 97\% of lost alignment without retraining and generalizes across unseen prompts, models, and quantizers.
Chat is not available.
Successful Page Load