Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
Abstract
Lay Summary
Training robots to act reliably in the real world usually requires many real-world videos, which are expensive and slow to collect. Computer simulations can generate robot videos much more cheaply, but simulated videos often look too clean, repetitive, or unrealistic, so robots trained on them may not work well in real environments. This paper introduces a way to turn simulated robot videos into realistic training videos while keeping the original task and robot motion unchanged. Our method first identifies what is happening in the simulated video, then describes the scene in natural language, changes the description to create more diverse environments, and finally generates a realistic version of the video that still follows the same action. To make this practical for large datasets, we also reduce unnecessary video generation and speed up the generation process. Experiments across several robot-learning benchmarks and a real robot show that the resulting videos help robots perform tasks more successfully and transfer better from simulation to the real world.