Because Flat Retrieval Is Not Enough: VideoGraph for Long Video Question Answering
Abstract
Long videos require systems to recover sparse evidence over time from both visual and spoken content. We present VideoGraph, a training-free method that builds a persistent multimodal graph memory once per video from transcript segments, visual clips, topics, and entities. We use long video question answering as an indirect probe of whether such a memory helps recover temporally grounded evidence. VideoGraph outperforms all compared baselines, including prior graph-based methods, on NExT-QA (77.62\%), EgoSchema (70.80\%), and Video-MME (69.11\%, medium split). The strongest supporting pattern is on NExT-QA, where gains are largest on temporal and causal questions, consistent with the intended role of the graph as a memory over temporally distributed and cross-modal evidence. Because the comparisons use different answer models and preprocessing pipelines, the results should be read as evidence that the approach is promising rather than as proof that graph memory alone explains the gains.