Representation Is Not Reliance: Concept-Based Causal Diagnostics of Shortcut Learning in Vision
Abstract
Vision models systematically exploit predictive but task-irrelevant visual cues as spurious shortcuts. This typically manifests as a severe performance drop when models generalize from in-domain to out-of-domain data. Although such shortcuts range from obvious artifacts to subtle signals, their visual form often remains elusive, even when their presence is statistically evident. To address this blind spot, we propose a diagnostic framework to expose the specific visual patterns that encode such shortcuts, paving the way for targeted removal. Specifically, we apply automated concept discovery to a shortcut-trained model to isolate candidate shortcut concepts and trace how localized visual patterns propagate across layers and become task-predictive. On a recently curated benchmark for investigating spurious correlations, these concepts lose predictive utility under shift despite being predictive under bias. Reusing the same exemplar-defined concepts in a model trained without the shortcut bias shows that these patterns remain separable and spatially coherent, and regain predictive alignment when the shortcut distribution is restored. Across depth, the same localized concepts remain traceable through most layers and weaken only in the final stage, where spatial detail is pooled before classification. Targeted interventions further show that the identified shortcuts are carried by localized visual regions and that weakening them changes predictive reliance, suggesting that shortcut behavior is governed by whether spurious cues are encoded, how they are organized, and when and where they drive prediction.