Visual Evidence Collapse as the Missing Signal in Adaptive Multimodal Retrieval: A Position
Göktuğ Aslanoğlu
Abstract
We take the position that adaptive retrieval in multimodal Retrieval-Augmented Generation (RAG) is held back by a structural miscalibration: the dominant classes of gating signals are blind to \emph{visual evidence collapse}, a failure mode in which reasoning VLMs silently abandon visual grounding while token entropy simultaneously \emph{decreases} making the collapse invisible to text-centric monitors. We propose the Visual Attention Trajectory (VAT), a gating signal that requires no auxiliary model or fine-tuning and monitors the windowed decay of attention mass on visual token regions during generation as a direct proxy for grounding quality. VAT is embedded in a three-action streaming routing framework (MR$^2$) for incremental multimodal QA; the broader framework requires a lightweight calibration step for its utility functions, but the VAT signal itself is training-free. We call on the community to adopt VAT as a benchmark signal for adaptive multimodal retrieval, and to establish QANTA-style streaming benchmarks that report retrieval rate, cost, and accuracy-efficiency trade-offs as first-class metrics.
Chat is not available.
Successful Page Load