Beyond Classification: Continual Learning for Multimodal Retrieval
Abstract
While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored. Existing work often treats continual retrieval as a byproduct of class-incremental learning (CIL), applying off-the-shelf methods within narrow evaluation schemes that obscure retrieval-specific failure modes and overestimate performance. To address this, we introduce a principled evaluation framework for continual multimodal retrieval spanning diverse visual domains, and systematically evaluate common approaches within this setting. Our empirical analysis shows that standard CIL methods fail to yield meaningful gains in this more realistic and challenging scenario. To tackle this problem, we propose Dynamic Adapter Routing (DAR), a novel approach based on prototypes, LoRA adapters and model merging, which outperforms existing methods by 8\%. We hope our work highlights the unique challenges of continual retrieval and encourages further research in this direction.