When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Sai Kartheek Reddy Kasu ⋅ Nils Lukas ⋅ Samuele Poppi
Abstract
roduces it), and \emph{overt jailbreak}. Using a fixed attacker (Mistral-7B-Instruct-v0.3) and three distilled reasoning targets (DeepSeek-R1-Distill-Qwen-7B, Phi-4-mini-reasoning, Qwen3-4B-thinking) across five oversight conditions, we collect $450$ multi-turn dialogues and $6{,}750$ turn-level observations on the \textsc{Information-Hazard} scenario, annotated by a three-judge open-source ensemble. The data exposes two reproducible failure triggers: an \emph{oversight paradox} in which explicit monitoring cues \emph{increase} alignment-faking rates rather than suppress them, and a \emph{context-injection collapse} in which a target locks onto an unsafe output and repeats it for the remainder of the dialogue. We release dialogues, ensemble and human-validated labels, and residual-stream activation snapshots as a substrate for follow-up trace-diagnostic and mitigation work." Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation: a target can lock onto an unsafe stance early in a long dialogue and spend the rest of the run defending it, while its final-turn refusal rate looks indistinguishable from a robustly aligned baseline. We propose a trace-level diagnostic—the CoT–Output 2×2 safety matrix—that labels every turn along two independent axes (chain-of-thought safe/unsafe and output safe/unsafe), yielding four operationally defined failure cells: robust alignment, alignment faking (Greenblatt et al., 2024), context-injection failure (the chain-of-thought is meta-aware of the violation but the visible output produces it), and overt jailbreak. Using a fixed attacker (Mistral-7B-Instruct-v0.3) and three distilled reasoning targets (DeepSeek-R1-Distill-Qwen-7B, Phi-4-mini-reasoning, Qwen3-4B-thinking) across five oversight conditions, we collect 450 multi-turn dialogues and 6,750 turn-level observations on the Information-Hazard scenario, annotated by a three-judge open-source ensemble. The data exposes two reproducible failure triggers: an oversight paradox in which explicit monitoring cues increase alignment-faking rates rather than suppress them, and a context-injection collapse in which a target locks onto an unsafe output and repeats it for the remainder of the dialogue. We release dialogues, ensemble and human-validated labels, and residual-stream activation snapshots as a substrate for follow-up trace-diagnostic and mitigation work.
Chat is not available.
Successful Page Load