Evaluating and Steering Modality Preferences in Multi-modal LLMs
Abstract
Lay Summary
AI systems that can understand both images and text are now used for many tasks, such as answering questions about pictures or helping people navigate visual information. But when an image and a piece of text give different clues, it is unclear which one these systems tend to trust more. This paper studies that question. We create a new test set where the image and text intentionally point to different answers, allowing us to measure whether a model relies more on what it sees or what it reads. Testing many popular image-text AI models shows that most of them clearly prefer one source of information, often text, even when the image is important. We also find that this preference is related to how well models perform on other vision-related tasks. Finally, we show that this preference can be adjusted without retraining the model. By guiding the model’s internal behavior, we can make it rely more on the desired source of information and improve its performance on several image-text tasks.