Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Abstract
Understanding modality interaction in multimodal large language models (MLLMs) is central to reliable deployment. We introduce Partial Information Decomposition (PID) as a decision-level framework that separates unique, redundant, and synergistic contributions of sensory and linguistic inputs, beyond representation alignment and outcome-based evaluation. Across vision--language benchmarks, PID reveals recurring modality-use profiles: reasoning and grounding-oriented tasks tend to exhibit high synergy, whereas expert and knowledge-oriented tasks show stronger language-unique reliance. These profiles generalize across model families and predict sensitivity to modality-level interventions. We further extend PID to tri-modal systems with Sensory PID, treating language as a control variable to decompose video--audio information gain. Applied to omni-modal models, Sensory PID reveals a sensory synergy bottleneck dominated by visual information even on audio--visual fusion tasks. Finally, PID-guided reweighting provides initial evidence for improving multimodal reasoning and grounding performance.
Lay Summary
When a multimodal AI model answers a question about an image, video, or sound, a correct answer does not tell us how it made the decision. Did it truly combine what it saw and heard, or did it rely mostly on language patterns and prior knowledge? This question matters as multimodal systems are increasingly used in settings where decisions need to be trusted and interpreted. Our work introduces an information-theoretic lens for opening this black box. We use Partial Information Decomposition to separate what each modality contributes uniquely, what is shared across modalities, and what only emerges when modalities are combined. This lets us distinguish models that genuinely integrate visual and textual evidence from those that rely more heavily on language priors. Across vision-language and omni-modal models, we find recurring patterns of modality use, including a visual dominance bottleneck in audio-visual tasks. We also show that these diagnostic signals can guide sample reweighting to improve multimodal reasoning and grounding. In short, our framework helps explain not only whether multimodal AI gets an answer right, but how it uses information to get there.