Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
Abstract
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework that explicitly anchors each reasoning step to specific visual evidence regions, enabling compositional and multi-step decision-making. Formally, Chain-of-Glimpse formulates video reasoning as a step-by-step process that incrementally builds spatially grounded traces around task-relevant visual objects, thereby mitigating over-reliance on saliency-driven cues. Specifically, Chain-of-Glimpse features a search-guided controller, optimized via reinforcement learning with a format reward that significantly incentivizes grounding capability, to iteratively ground visual evidence regions and form reliable reasoning trajectories, yielding accurate and interpretable multi-step decisions. Extensive evaluations across two categories of video reasoning benchmarks, including general video reasoning benchmarks such as NExTQA, Video-Holmes, CG-Bench-Reasoning, and VRBench, and grounded video reasoning benchmarks such as NExT-GQA, demonstrate that Chain-of-Glimpse consistently improves performance while exhibiting strong robustness and generalization across diverse video reasoning tasks.
Lay Summary
Video understanding systems must identify and reason over task-relevant objects across frames, but existing object-agnostic methods often fail when objects change appearance, position, or context over time. This limitation makes video reasoning vulnerable to saliency-driven shortcuts and weakens both accuracy and interpretability. To address this problem, we propose Chain-of-Glimpse, a search-guided progressive object-grounded reasoning framework. Rather than reasoning over videos holistically, Chain-of-Glimpse decomposes video understanding into step-by-step reasoning traces, where each step is explicitly anchored to specific visual evidence regions. A search-guided controller, optimized with reinforcement learning and a grounding-oriented format reward, iteratively selects task-relevant regions and constructs reliable multi-step reasoning trajectories. This design enables compositional decision-making while reducing over-reliance on visually salient but irrelevant cues. Experiments on general video reasoning benchmarks, including NExTQA, Video-Holmes, CG-Bench-Reasoning, and VRBench, as well as grounded video reasoning benchmarks such as NExT-GQA, show consistent performance gains. These results indicate that Chain-of-Glimpse improves not only accuracy but also robustness and generalization across diverse video reasoning tasks. More importantly, the framework makes video reasoning more interpretable by linking decisions to concrete visual evidence.