ExpAlign: Expectation-Guided Vision–Language Alignment for Open-Vocabulary Grounding
Abstract
Lay Summary
Artificial intelligence models often struggle to accurately find and outline specific objects in images based on open-ended, everyday language descriptions. Existing methods either produce blurry, imprecise results by compressing entire sentences into a single concept, or they demand massive computing power and tedious, manual human annotations to map out every single word. To bridge this gap, researchers developed a lightweight framework called ExpAlign to optimize word-to-image matching. Without relying on expensive manual data labeling, ExpAlign uses a smart mathematical mechanism to automatically calculate how strongly different image regions relate to specific words. It also features a unique geometric consistency technique that ensures the model locks onto coherent object shapes rather than fragmented boundaries. Across multiple benchmarks, ExpAlign demonstrates an exceptional ability to spot and precisely outline rare or unfamiliar objects that it was never explicitly trained to recognize. It substantially outperforms dominant, heavier state-of-the-art vision models on long-tail categories. Crucially, the module delivers these high-precision maps while remaining incredibly fast, lightweight, and efficient for practical deployment.