Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence
Abstract
The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question–answer (QA) pairs from a limited number of manually annotated datasets, rather than systematically annotating new large-scale 3D scenes from raw web data. As a result, their scalability is severely constrained, and model performance is further hindered by domain gaps inherent in these narrowly curated datasets. In this work, we propose \textbf{Holi-Spatial}, the first fully automated, large-scale, spatially-aware multimodal dataset, constructed from raw video inputs without human intervention, using the proposed data curation pipeline. Holi-Spatial supports multi-level spatial supervision, ranging from geometrically accurate 3D Gaussian Splatting (3DGS) reconstructions with rendered depth maps to object-level and relational semantic annotations, together with corresponding spatial Question–Answer (QA) pairs. Following a principled and systematic pipeline, we further construct \textbf{Holi-Spatial-4M}, the first large-scale, high-quality 3D semantic dataset, containing 12K optimized 3DGS scenes, 1.3M 2D masks, 320K 3D bounding boxes, 320K instance captions, 1.2M 3D grounding instances, and 1.2M spatial QA pairs spanning diverse geometric, relational, and semantic reasoning tasks. Holi-Spatial demonstrates exceptional performance in data curation quality, significantly outperforming existing feed-forward and per-scene optimized methods on datasets such as ScanNet, ScanNet++, and DL3DV. Furthermore, fine-tuning Vision-Language Models (VLMs) on spatial reasoning tasks using this dataset has also led to substantial improvements in model performance.
Lay Summary
AI systems that understand the 3D world need large amounts of detailed training data. For example, they need to know where objects are, how they are arranged, which objects are close to each other, and how a scene looks from different viewpoints. However, existing datasets for this kind of spatial understanding are usually built from a small number of manually labeled scenes. This makes them expensive to create, difficult to scale, and less representative of the variety found in real-world environments. In this work, we introduce Holi-Spatial, a fully automated system for building large-scale 3D spatial understanding data directly from raw videos, without manual annotation. Given videos, our system reconstructs detailed 3D scenes, identifies objects, describes their spatial relationships, and creates questions and answers that can be used to train and evaluate AI models. Using this system, we build Holi-Spatial-4M, a large dataset containing thousands of reconstructed 3D scenes and millions of spatial annotations and question-answer examples. Our experiments show that this dataset provides high-quality spatial information and helps vision-language AI models perform better on tasks that require understanding 3D scenes. We hope Holi-Spatial will support future AI systems that can better understand, reason about, and interact with the physical world.