SeeTraceAct: Grounding Vision-Language-Action Models with One-Shot Video Demonstrations via Visual Latent Planning
Abstract
Vision-language-action models (VLAs) have shown strong promise as general-purpose robot policies. However, adapting them to new tasks typically requires costly task-specific teleoperation data. In contrast, human demonstration videos are easier to collect, making them an attractive source of supervision for task adaptation. In this work, we study one-shot demo-conditioned VLAs, where a robot is conditioned on a single demonstration video of the target task. We find that existing end-to-end approaches often treat the demonstration primarily as a task identifier, rather than as an executable visual plan for spatially precise action generation. To address this limitation, we propose SeeTraceAct, a demo-conditioned VLA framework that learns a visual latent plan from demonstrations via auxiliary future visual trace supervision. In our experiments on RoboCasa benchmarks, SeeTraceAct significantly outperforms strong baselines, with the largest gains on tasks that demand precise spatial interaction, showing that our approach is effective for grounding VLAs on one-shot demonstration videos.