RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization
Abstract
Lay Summary
Robots are still difficult to train for everyday tasks because most systems only work well in the specific settings and on the specific robot hardware they were trained on. This makes it costly to adapt a robot to new homes, objects, instructions, or robot arms. In this work, we introduce RDT2, a general-purpose robot model designed to follow language instructions and control different robots without needing new training for every new situation. To build it, we collected over 10,000 hours of real-world demonstrations using a portable hand-held device, allowing people to teach many household manipulation skills in diverse environments. We also design a training process that helps the model connect what it sees and hears with precise robot movements while still running fast enough for real-time control. Experiments show that RDT2 can handle new objects, scenes, instructions, and robot platforms, and performs strongly on challenging tasks such as cloth folding, table bussing, button pressing, and table tennis. This suggests a practical path toward more adaptable robot assistants that can be deployed more broadly with less task-specific data collection.