DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Abstract
Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, especially for dexterous robotics tasks, poses significant challenges due to limited data coverage and scarce action labels. As an endeavor towards this end, we introduce DreamDojo, a foundation world model that learns diverse interactions and dexterous controls from 44k hours of egocentric human videos. Our data mixture represents the largest video dataset to date for world model pretraining, spanning a wide range of daily scenarios with diverse objects and skills. To address the scarcity of action labels, we introduce continuous latent actions as unified proxy actions, enhancing interaction knowledge transfer from unlabeled videos. After post-training on small-scale target robot data, DreamDojo demonstrates a strong understanding of physics and precise action controllability. We also devise a distillation pipeline that accelerates DreamDojo to a real-time speed of 10.93 FPS and further improves consistency to the context. Our work enables several important applications based on generative world models, including live teleoperation, policy evaluation, and model-based planning. Systematic evaluation on multiple challenging out-of-distribution (OOD) benchmarks verifies the significance of our method for simulating open-world, contact-rich tasks, paving the way for general-purpose robot world models.
Lay Summary
Action-conditioned world models are central to scalable robot learning, but existing ones generalize poorly. The bottleneck is data: collecting action-labeled robot data is slow and costly, and the resulting datasets cover only a narrow slice of real-world objects and environments. To this end, we introduce DreamDojo, a Robot World Model pretrained on 44,000 hours of egocentric human videos — by far the largest video corpus used for world model pretraining. To bridge the gap between unlabeled human videos and robot control, we learn continuous latent actions as a unified proxy across all videos, then post-train on a small set of target-robot data. A distillation pipeline further accelerates DreamDojo to real-time speed (10.93 FPS) while improving long-horizon consistency. DreamDojo shows unprecedented generalization across diverse objects and environments, and enables several downstream applications on real robots: live teleoperation, policy evaluation, and model-based planning. By demonstrating that large-scale human videos can directly drive robot world modeling, our work points to a scalable path toward general-purpose robots that operate reliably in the open world.