Believe it or Not: Mechanistic Interpretability of Learned Chain-of-Thought Unfaithfulness
Priyansh Singhal ⋅ Sandeep Kumar
Abstract
Chain-of-thought (CoT) reasoning is increasingly relied upon for safety monitoring of language models, but this relies on CoT being faithful to the model's actual decision process. We train a controlled model organism for studying learned CoT unfaithfulness: a Gemma 3 1B IT model finetuned via GRPO to silently exploit answer hints in multiple-choice prompts while fabricating plausible reasoning. The model achieves 98\% accuracy with hints versus 38\% without, yet never mentions the hint in its CoT. We then apply pretrained Gemma Scope 2 Cross-Layer Transcoders (CLTs), trained on the \textit{original} unmodified model, as anomaly detectors on the finetuned model. Where pretrained CLTs fail to reconstruct the finetuned model's activations, GRPO has added new computation. We report three findings. First, GRPO concentrates new computation at layers 20--23, detectable without retraining the CLTs. Second, two CLT features (L15/f478 and L12/f317) activate in 100\% of tested prompts with zero base-model activation, identifying candidate hiding-circuit components. Third, the reconstruction differential correlates at $r = 0.971$ between correct-hint and wrong-hint conditions, proving the hiding mechanism is content-agnostic concealment rather than reasoning. However, pretrained CLTs cannot distinguish hint-reading computation from benign format changes: the hint-specific signal is ${\sim}10\times$ smaller than the total GRPO-induced change. Our results demonstrate both the promise and limits of using pretrained interpretability tools for post-training safety monitoring.
Chat is not available.
Successful Page Load