VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation
Abstract
Lay Summary
Videos often contain many moving objects, and people may refer to a target using complex descriptions involving actions, timing, or relationships. Existing video segmentation methods usually rely on a fixed set of sampled frames, so they can miss the brief but crucial visual evidence needed to identify the correct object. We tackle this with VideoSEG-O3, a system that analyzes videos through multi-step visual exploration. It first forms a broad understanding of the video, then actively selects important time intervals and key frames to inspect more closely before producing the final object mask. We also design a training strategy that connects the model’s reasoning process with pixel-level mask quality, helping it learn both where to look and how to segment the target accurately. This research makes language-guided video object segmentation more reliable, especially for long videos and descriptions that require reasoning about motion or events. It can support applications such as video editing, robotics, surveillance analysis, and assistive visual tools, where systems need to find the right object from natural language instructions.