G$^2$TAM: Geometry Grounded Track Anything Model
Abstract
Lay Summary
AI systems are increasingly able to segment and track objects in videos, but they often rely on how objects look from frame to frame. This can fail when the camera moves a lot, when an object is hidden for a long time, or when the same object looks very different from another viewpoint. In this work, we ask whether 3D spatial understanding can help an AI system remember and track objects more reliably. We introduce G2TAM, a model that jointly reconstructs scene geometry and segments a prompted object across images or videos. Instead of storing only visual appearance as memory, the model uses a geometry-aware representation as a more stable form of spatial memory. A user can specify the target object with a point, box, or text description, and the model predicts consistent object masks across different views and time steps. This matters because reliable object tracking is important for applications such as robotics, augmented reality, and interactive scene understanding. By connecting segmentation with 3D geometry, our work moves toward AI systems that understand not only what objects look like, but also where they are in the surrounding space.