XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
Abstract
Lay Summary
XR-1: A General Robot Model That Connects Seeing and Acting Across Different Robots AI systems are now very good at chatting, writing, and creating images, but robots still struggle in the physical world. A robot may see a complex scene, yet fail to turn that visual information into precise fingertip movements, and it often cannot learn a new skill simply by watching a video as humans do. To address this challenge, we developed XR-1, a general robot model designed to work across different robot bodies. At its core is UVMC, a shared “visual-action language” that acts like a translator between what a robot sees and how it moves. UVMC converts both visual motion and physical actions into a unified digital representation, allowing XR-1 to bridge differences in robot hardware and even learn useful action patterns from human demonstration videos. We tested XR-1 in 14,000 real-world trials across 6 robot platforms and more than 120 tasks, including bimanual pouring and precise door opening. XR-1 outperformed leading existing models and remained robust in unfamiliar cluttered scenes and changing lighting conditions. These results bring us closer to building general-purpose robot assistants that can adapt to many tasks, environments, and bodies.