Hallucination-Induction, Not Calibration: When Multi-Feature SAE Steering Looks Like It Works
Caio S Vicentino
Abstract
We replicate Ferrando et al.'s (2024) entity-recognition SAE feature pattern on a 27B reasoning-tuned model (Qwen3.6-27B) using a paper-grade Top-$K$ Sparse Autoencoder, then test whether the feature is causally usable for hallucination calibration. The single best SAE latent at layer 31 reaches AUROC $0.814$ (95% bootstrap CI $[0.731, 0.898]$) on cross-type known-vs-unknown classification, with the entity-recognition signal peaking in the mid-stack as Ferrando reported on Gemma-2-2B-IT (peak around layer 9 in their 26-layer model). The SAE latent is statistically indistinguishable from a layer-32 L2 logistic regression probe ($0.887$ $[0.823, 0.941]$) and a diff-of-means probe ($0.859$): the contribution is interpretability, not raw classification performance. Single-feature steering produces a null effect on refusal calibration. Multi-feature top-$K$ ablation at small $K{=}200$ ($0.3$% of the dictionary) produces a $4$–$8\sigma$ effect against a random-$K$ null — in contrast with prior large-$K$ work where the random control nullifies the signal (Al-Qurashi 2025) — but a Claude-as-judge correctness audit reveals the effect is *induced confabulation*: known-entity incorrect-answer rate rises from $62$% to $77$% while correct-answer rate drops from $8$% to $0$%. Statistical significance against random-$K$ is necessary but not sufficient evidence for a calibration mechanism: the intervention identifies a hallucination-induction circuit, not a calibration knob.
Chat is not available.
Successful Page Load