ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation
Abstract
Incremental Learning (IL) for Open-ended Image-to-Text Generation (OpenITG) enables models to continuously generate accurate, contextually relevant text for new images while preserving previously acquired knowledge. Unlike prior studies, this paper addresses a more practical scenario in which the predominant category of visual data shifts over time as environments evolve. In this context, we introduce a new notion of continual alignment, which incrementally adapts the alignment module within pre-trained VLMs to preserve high-quality cross-modal representations. Based on this idea, we propose Efficient Continual Alignment (ECA), a novel exemplar-free IL approach for OpenITG. The key challenge is enabling the model to acquire new, task-specific features while minimizing interference with the established alignment without accessing raw data from previous tasks. To address this, ECA employs three core mechanisms: a Mixture of Query (MoQ) module that adapts task-specific query tokens, a Fisher Dynamic Expansion (FeDEx) that dynamically expands model structure based on a Fisher Information Matrix (FIM)-based metric, and an embedding dictionary with Dictionary Replay (DR) to retain past knowledge. To evaluate ECA's performance, we construct four new IL OpenITG benchmarks that better reflect real-world scenarios. Experimental results demonstrate that ECA significantly mitigates catastrophic forgetting and improves IL performance compared to baseline methods. Code and benchmarks are available at https://github.com/Snowball0823/ECA.
Lay Summary
AI systems that write descriptions of images or answer questions about images often need to keep learning after they are deployed. In real use, the images they receive over time may shift from indoor scenes to vehicles, food, or other subjects, while earlier subjects can still appear in the background. Teaching the system with each new group of images can help it adapt, but it can also make it worse at handling images it learned before, especially when old images cannot be kept. This paper studies how to help such systems learn from changing images without losing earlier abilities. Our method keeps most of the original system fixed and updates only the small part that connects image understanding with text generation. It also keeps a compact record of past visual patterns instead of storing old images. We create four test settings that better reflect changing image content over time. The results show that our method helps the system remember earlier image types while learning new ones.