Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
Abstract
Video-Language Models (VidLMs) achieve strong benchmark scores, yet it remains unclear whether they truly encode and use visual evidence when generating answers. Aggregate accuracy alone cannot distinguish whether failures arise because visual signals were never encoded or because they were later overridden by model priors. We introduce REVEAL, a controlled diagnostic framework for linking VidLM behavior to internal failure mechanisms. REVEAL contains five probes: camera-motion sensitivity, cross-frame integration, video sycophancy, language-only shortcuts, and temporal expectation bias. Together, these probes test whether models encode basic video signals, integrate evidence across frames, and preserve visual evidence against linguistic and prior-driven biases. Across 11 VidLMs, we find systematic failures along both pathways. We further conduct mechanistic analyses to localize where visual evidence is encoded, ignored, or suppressed across the model pipeline. More broadly, the encoding-versus-override distinction extends beyond VidLMs and provides a general framework for diagnosing evidence loss in multimodal systems.