M-IDoL: Information Decomposition for Modality-Specific and Diverse Representation Learning in Medical Foundation Model
Abstract
Medical foundation models (MFMs) aim to learn universal representations from multimodal medical images that can generalize effectively to diverse downstream clinical tasks. However, most existing MFMs suffer from information ambiguity that blends multimodal representations in a single embedding space, leading to the degradation of modality specificity and diversity. In this paper, we propose M-IDoL, a self-supervised MFM that introduces Information Decomposition for multimodal representation Learning via two objectives: i) maximizing inter-modality entropy by dispersing multimodal representations into separable Mixture-of-Experts (MoE) subspaces to achieve representation specificity across modalities; and ii) minimizing intra-modality uncertainty by performing fine-grained semantic discrimination within each MoE subspace to enrich representation diversity per modality. By pre-training on 1.15 million medical images, M-IDoL i) delivers superior generalization across 21 downstream clinical tasks, outperforming 20 foundation models on five imaging modalities (e.g., X-ray, fundus, OCT, dermoscopy and pathology), and ii) learns modality-specific and diverse representations, showing clearer separation of feature clusters across modalities and finer-grained feature discrimination within each modality.
Lay Summary
Medical images generally come from many different sources, such as X-ray, fundus, OCT, dermoscopy and pathology. These imaging modalities contain different types of information, making it difficult for a single foundation model to learn effective visual representations from all of them. We propose M-IDoL, a self-supervised learning model that enhances visual representations by capturing specific and diverse information from different medical imaging modalities. Instead of mixing all image information together, M-IDoL separates modality-specific features and reduces confusion between image types. This enables the model to better understand what is unique to each modality while still learning broadly useful medical image representations. Across 21 downstream medical imaging datasets, M-IDoL consistently performs better than existing state-of-the-art medical foundation models. These results show that M-IDoL can generalize well across different clinical tasks and imaging domains, making it a promising approach for robust medical image analysis.