Lightweight Alignment of Unimodal Foundation Models for Metabolite Identification
Abstract
A central challenge in building multimodal foundation models for the life sciences is the imbalance between abundant unimodal data and scarce paired observations, which limits the scalability of joint multimodal pretraining. We investigate an alternative approach based on aligning pretrained unimodal models. Focusing on metabolite identification, we introduce MSAlign, which maps a molecular transformer (ChemBERTa) and a mass spectra transformer (DreaMS) into a shared embedding space. Despite its simplicity, MSAlign substantially outperforms prior methods across benchmarks, setting a new state-of-the-art in retrieval performance. These results suggest that aligning unimodal foundation models offers an effective route to multimodal learning in biological settings where paired data remain limited.