The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence
Abstract
While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment--uniformity decomposition. We develop a measure-theoretic framework in which representation measures evolve on a fixed embedding manifold. In the large-batch limit, we prove value and gradient consistency, linking the stochastic objective to explicit deterministic energy landscapes and revealing a geometric bifurcation between unimodal and symmetric multimodal regimes. In the unimodal case, the intrinsic energy is strictly convex and admits a unique Gibbs equilibrium, showing that entropy acts as a tie-breaker within the aligned basin. In the multimodal case, the intrinsic geometry becomes cross-coupled and contains a persistent negative symmetric divergence term: each modality's marginal reshapes the effective landscape of the other, allowing strong pairwise alignment to coexist with a persistent modality gap. Controlled synthetic experiments and analyses of pretrained CLIP representations support these predictions. Overall, our results shift the analytical lens from pointwise discrimination to population geometry, showing that pairwise alignment alone is insufficient to control cross-modal marginal structure.
Lay Summary
Modern image–text AI systems often learn by pulling matching image–caption pairs together and pushing unrelated pairs apart. However, even when individual images and captions are well matched, the overall image and text representation clouds can remain separated. This paper studies why this "modality gap" can arise. We show that InfoNCE, a widely used contrastive learning objective, shapes not only pairwise matches but also the geometry of the full representation populations. In single-modality learning, this geometry tends to be cohesive. In multimodal learning, however, the image and text populations can interact in a way that sustains separation, even under strong pairwise alignment. Experiments on synthetic data and pretrained large vision-language models support this explanation. Our results suggest that future multimodal systems should be evaluated and trained not only for retrieval accuracy, but also for how well their full image and text representation distributions align.