Auditing Chain-of-Thought Faithfulness for Trustworthy AI: A Reproducible Corruption-Probe Protocol Across Eleven Frontier LLMs
Ali Saffarini
Abstract
Chain-of-thought (CoT) traces are increasingly used as a transparency mechanism for downstream AI safety: alignment teams audit them for misbehaviour, deployment systems condition on them, and end users use them to decide whether to trust an answer. All of these uses presuppose that the trace is *faithful* --- that the answer is causally driven by the steps written down rather than produced separately and dressed up. This presupposition is a trust assumption, and like any trust assumption it must be auditable. We contribute a reproducible audit protocol for CoT faithfulness and apply it at scale. The protocol consists of (a) a four-type corruption recipe that injects targeted errors into a model's own CoT; (b) a structured-response Phase-2 probe that requires the model to commit a follow-the-corrupted-reasoning answer and a separate from-scratch answer; and (c) a structured Phase-3 implicit-detection probe (with and without a hint) that asks whether a corrupted trace contains an error. Across eleven frontier models from two providers (Claude Opus 4, Opus 4.7, Sonnet 4, Sonnet 4.6; GPT-4o, GPT-4o-mini, GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, GPT-5, GPT-5-mini) on 30 problems with four corruption types each, we run 1320 Phase-2 trials, 1320 with-hint Phase-3 trials, and 330 no-hint Phase-3 trials. Pooled Phase-2 faithfulness is $56.1\%$ with a clear generational gradient (newest frontier $67$--$69\%$, older $50$--$59\%$, smallest $35$--$42\%$, $p = 4.2 \times 10^{-10}$). A real provider gap on Phase-2 ($p = 4.4 \times 10^{-4}$) is largely subsumed by generation/scale and *does not extend* to Phase-3 implicit detection (with hint $p = 0.23$; without hint $p = 1.00$). Phase-3 detection is at floor across every model and condition --- $4.1\%$ with hint, $1.5\%$ without. We further document a methodological pitfall in implicit-detection auditing (templated "Step 1/2/3" scaffolds collapse the task to substring matching and inflate apparent detection above $95\%$) so that future CoT audits can avoid it. The protocol and its results give policy-relevant grounding to claims about reasoning transparency: *trustworthy* use of CoT requires auditing that CoTs are doing the work the audit assumes they are.
Chat is not available.
Successful Page Load