Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction
Abstract
Recovering scene-consistent 4D crowd motion from monocular video in large-scale scenes remains challenging due to severe depth ambiguity and complex scene geometry. Existing monocular crowd reconstruction methods typically rely on single-plane assumptions, leading to unreliable metric scale and spatial drift under complex terrain. We propose Crowd4D, the first scene-aware 4D crowd reconstruction framework that jointly optimizes the crowd and scene from a monocular RGB video in large-scale scenes. Crowd4D explicitly incorporates scene geometry and ensures consistency across image and scene spaces via a multi-stage optimization strategy. A key bottleneck of this task lies in accurate human–scene alignment, particularly in scale and position. However, human and scene reconstructions are typically decoupled. To address this, we introduce the Human–Scene Interaction Proxy (HSIP) as an intermediate representation, derived from Scene Interaction Point Clouds and a Scene Interaction Surface (SIPC&SIS), which encode explicit scene-aware geometric priors and redefine the optimization space for large-scale monocular 4D crowd reconstruction. To further improve temporal stability under occlusions, we introduce Crowd Structural Coherence Regularization (CSCR), which leverages HSIP-based spatial priors to impose soft temporal consistency on pairwise relative displacements and directions within local crowd neighborhoods. Extensive experiments demonstrate that Crowd4D consistently outperforms existing state-of-the-art methods and enables robust monocular 4D crowd reconstruction in complex, large-scale real-world scenes. Project page is available at https://cic.tju.edu.cn/faculty/likun/projects/Crowd4D.
Lay Summary
Understanding how large groups of people move through real environments is important for analyzing public spaces, supporting crowd management, and creating realistic digital content. However, recovering this motion from a single ordinary video is difficult, especially when people are far away, partially blocked from view, or moving over slopes, stairs, and uneven ground. We present Crowd4D, a method that reconstructs the three-dimensional motion of many people over time while also considering the shape of the surrounding scene. Unlike previous approaches that often assume people move on a flat surface, our method uses the estimated environment to better determine where people are and how they move through it. It also improves stability when individuals are temporarily hidden by using the spatial relationships among nearby people. Experiments show that Crowd4D produces more accurate and stable reconstructions than previous methods in large and complex scenes, providing a practical way to understand real-world crowd motion from widely available video recordings.