UniFast-HGR: Scalable and Efficient Maximal Correlation for Multimodal Models
Abstract
Lay Summary
Many AI systems need to learn from several kinds of information at the same time, such as images, text, audio, or satellite data. A key challenge is to find what these sources truly share, while ignoring noise or details that only appear in one source. Existing methods that measure such shared information can be powerful, but they often become too expensive or unstable when used in modern large models. This paper introduces a simpler training method that keeps the useful idea of measuring shared dependence, but avoids the costly calculations that usually make it hard to scale. It also removes unhelpful self-comparisons that can dominate the training signal. A faster version further reduces memory use for large batches. Experiments on image recognition, text-image and video-text retrieval, remote-sensing segmentation, and emotion recognition show that the method improves accuracy while remaining efficient for very large feature dimensions.