VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
Abstract
Lay Summary
Understanding a video is more than recognizing objects or actions. To truly understand what is happening, an AI system must know both when an event happens and where it occurs on the screen. Humans do this naturally, but existing methods process time and space separately, making them struggle in complex real-world scenes. In this work, we introduce VideoLoom, a Video LLM designed to connect spatial and temporal understanding within a single framework. We also build a new dataset that teaches the model to jointly track events across both time and space, as well as a new benchmark for evaluating this capability in a comprehensive way. Our approach outperforms previous methods on video tasks that require locating actions and objects in time, in space, or jointly across both. By helping AI systems interpret videos in a more unified and human-like way, our work may support future applications such as video assistants, robotics, and intelligent surveillance systems.