Causal Sufficiency Without Semantic Alignment: How Causal Subspaces Can Masquerade as Semantic Concepts
Abstract
Mechanistic interpretability treats a causally active subspace as evidence that a model encodes the corresponding concept. We apply Distributed Alignment Search to arithmetic verification across three instruction-tuned LLMs and find a causally sufficient ``arithmetic truth'' subspace, which is capable of steering incorrect predictions towards correct ones. However, on held-out samples, we find that a probe trained to decode the ground truth performs at chance, while a probe trained to decode the model's prediction achieves perfect accuracy. This indicates that the subspace encodes prediction bias, not truth. CoT prompting repairs the model precisely by routing around the causal subspace. We conclude that a localised subspace can be causally sufficient but semantically misaligned with the ground truth concept it is meant to capture.