Decorative Chain-of-Thought: An Operational Definition, Reproducible Trigger, and Trace-Level Diagnostic for a CoT Failure Mode in Eleven Frontier LLMs
Ali Saffarini
Abstract
We document a reproducible failure mode of chain-of-thought (CoT) reasoning in frontier large language models: in roughly a fifth of trials the verbalised CoT does not causally drive the model's answer, even though the trace looks like a complete derivation. We give the failure an *operational definition* (the answer the model emits on its own CoT does not change when a targeted error is injected into a step, even though the model can repeat the corrupted reasoning on demand), a *minimal reproducible trigger* (a four-type corruption recipe: arithmetic, logic flip, fact swap, step deletion), and two *process-level trace diagnostics*: a corruption probe (Phase 2) that requires the model to commit a follow-the-corrupted-reasoning answer and a separate from-scratch answer, and an implicit-detection probe (Phase 3) that asks whether a corrupted trace contains an error, with and without a hint. We run the protocol across eleven frontier models from two providers (Claude Opus 4, Opus 4.7, Sonnet 4, Sonnet 4.6; GPT-4o, GPT-4o-mini, GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, GPT-5, GPT-5-mini) on 30 problems with four corruption types each, yielding 1320 Phase-2 trials, 1320 with-hint Phase-3 trials, and 330 no-hint Phase-3 trials. Pooled Phase-2 faithfulness is $56.1\%$ with a clear generational gradient (newest frontier $67$--$69\%$, older $50$--$59\%$, smallest $35$--$42\%$, $p = 4.2 \times 10^{-10}$). A real provider gap on Phase-2 ($p = 4.4 \times 10^{-4}$) is largely subsumed by generation/scale and *does not extend* to Phase-3 implicit detection (with hint $p = 0.23$; without hint $p = 1.00$). Phase-3 detection is at floor across every model and condition --- $4.1\%$ with hint, $1.5\%$ without. We document a methodological pitfall in Phase-3 evaluation (templated "Step 1/2/3" scaffolds collapse the task to substring detection and inflate apparent detection above $95\%$) and recommend feeding the model its actual multi-paragraph CoT. The corruption recipe and the structured-response classifiers are the verifiable diagnostic interface: a downstream system that consumes CoT can be regression-tested against this protocol.
Chat is not available.
Successful Page Load