SIGMMA: Hierarchical Graph-Based Multi-Scale Multi-modal Contrastive Alignment of Histopathology Image and Spatial Transcriptome
Abstract
Recent advances in computational pathology have leveraged vision–language models to learn joint representations of Hematoxylin and Eosin (H&E) images with spatial transcriptomic (ST) profiles, but existing approaches typically align H&E tiles and ST profiles at a single scale, overlooking fine-grained cellular structures and their spatial organization. We propose \textsc{Sigmma}, a multi-modal contrastive alignment framework for learning hierarchical H&E-ST representations. By enforcing multi-scale contrastive alignment, \textsc{Sigmma} ensures coherent representations across modalities, while a graph-based modeling of cell interactions integrates both inter- and intra-subgraph relationships to capture cellular organization from fine to coarse scales. Across datasets, \textsc{Sigmma} consistently improves gene-expression prediction and cross-modal retrieval performance, and its learned multi-scale embeddings recover tumor microenvironments and immune-exclusion programs in pancreatic cancer.