CCDiff: Inverse Canonical Correlation Analysis for Discovering Visual Differences in Natural Language
Neelesh Bisht ⋅ Xingjian Li ⋅ Zihan Li ⋅ Bo Jiang ⋅ RUNMIN JIANG ⋅ Mostofa Rafid Uddin ⋅ Yang Liu ⋅ Min Xu
Abstract
Set-level visual difference discovery is increasingly important for dataset auditing and for understanding model behavior under distribution shift, yet manually inspecting thousands of images is impractical. This motivates the emerging task of *set difference captioning*, i.e., describing concepts that are more often true for one image set than another. Different from existing approaches which heavily rely on large language models (LLMs), incurring substantial computational cost in terms of time and tokens, we introduce **CCDiff**, a lightweight, statistically grounded, and training-free framework for set difference captioning that significantly reduces LLM dependence during core difference discovery, achieving over $2\times$ speedup and reducing token usage to zero. CCDiff operates in three stages: constructing a shared pool of candidate concepts from domain vocabularies or lightweight captioning models, filtering candidates to obtain set-specific concept pools, and performing *inverse* canonical correlation analysis (CCA) across sets to identify low-correlation directions that correspond to set differences. Candidate concepts are then ranked using a CCDiff score that balances inverse correlation with within-set representativeness. We evaluate CCDiff on VisDiffBench and further validate it on additional domains including MetaShift and MIMIC-CXR. Across benchmarks and real-world settings, CCDiff matches or exceeds prior LLM-based approaches while being substantially more efficient and scalable.
Chat is not available.
Successful Page Load