Compositional by Design: Background-Invariant Representations via Linear Additivity in VLMs
Abstract
Vision-language models (VLMs), such as CLIP and SigLIP 2, are widely deployed, yet they often fail to systematically generalize out-of-domain, relying on statistical shortcuts rather than true compositional understanding. A prominent and practically critical failure mode is background-based spurious correlations, where models improperly entangle reusable components—foreground objects and their backgrounds. In this paper, we systematically quantify how VLMs internally represent compositionality, specifically examining the linear additivity of foreground and background concepts. Leveraging these internal representational structures, we develop a new mitigation method, Background-invariant Anchor Pre-training (BAP). Our method explicitly isolates compositional features to achieve state-of-the-art worst-group accuracy exceeding 90% on Waterbirds under perfect spurious correlation (no minority-group examples in the training data). BAP demonstrates how enforcing modular, compositional structures can drive robust, out-of-domain generalization across benchmarks. Highly practical, it relies exclusively on synthetic data and exhibits strong sim-to-real transfer, paving the way for safer deployment in real-world scenarios