Prior Dominance in Audio-Visual LLMs: When Generative Models Memorize Over Reasoning Under Cross-modal Conflict
Abstract
Audio hallucinations remain a critical failure mode in Audio-Visual LLMs, where models sub- stitute memorized distributional priors for actual audio grounding when forced to process conflict- ing cross-modal inputs. Using the logit lens and head shutoff ablations on VideoLLaMA 2-7B-AV, we track the internal trajectory of cross-modal rea- soning. We find that model commitment concen- trates at layer 25.5 ± 1 across all configurations, regardless of alignment intervention. Crucially, 21 conflict-resolution heads cluster at layers 15-18, structurally upstream of this commitment point, indicating the model detects conflict early but overwrites it during generation. Behaviorally, all three fine-tuned configurations and off-the-shelf InternVideo2 collapse to near-chance on conflict examples, suffering a 32.3% accuracy drop and 17.3% instruction failure rate while shifting out- put priors. These findings identify late-layer prior commitment, not early-stage temporal alignment, as the key intervention target.