Detecting Hidden Chain-of-Thought in Large Language Models With Linguistic, Behavioral, and Mechanistic Indicators
Armaan Singh ⋅ Ryan T Le ⋅ Jasmine Kaur ⋅ Edward L Lip ⋅ Kiran Nijjer ⋅ Adnan Ahmed ⋅ Vasu Sharma
Abstract
Large language models often answer complex reasoning questions correctly without revealing intermediate steps, raising the question of whether they perform latent reasoning or pattern completion. Existing detection methods rely on surface-level linguistic cues that prior work has shown can be unfaithful. We propose the Hidden CoT Detection Score (HCDS), a comparative behavioral and mechanistic signal that quantifies whether a model's neutral-prompt behavior aligns more closely with its explicit-CoT or with its explicit no-CoT behavior. On Qwen3-4B Instruct, HCDS is robustly positive on both GSM8K and StrategyQA ($+1.79$ to $+2.38$, $p < 10^{-10}$) --- our cleanest hidden-CoT detection result. On Qwen3-4B Thinking, HCDS is also positive ($+0.33$ to $+0.60$, $p < 0.05$) but the result is better read as evidence of prompt-invariance of reasoning traces: under explicit no-CoT directives, Thinking emits a mean of ${\sim}575$ reasoning tokens vs Instruct's ${\sim}24$, so the no-CoT pole is not a clean answer-only baseline for that model. Anchor-suppression analysis further shows that the Thinking model distributes its causal load across many reasoning steps, suggesting that hidden chain-of-thought in reasoning-tuned models is more deeply integrated, more prompt-invariant, and more diffuse than in instruction-tuned ones.
Chat is not available.
Successful Page Load