LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries
Abstract
Lay Summary
Robots are increasingly trained to follow human instructions, such as picking up a specific object or placing it in a certain location. However, many robot training datasets unintentionally make the task too easy to guess from the image alone. For example, if a robot usually sees the same scene whenever it is asked to perform a certain task, it may learn to act based only on visual patterns rather than actually understanding the instruction. This can make the robot appear successful during training, but fail when the same scene could require different actions. We study this problem and propose LangForce, a training method that encourages robots to keep using language when choosing actions. LangForce compares what the robot would infer from the image alone with what it should do when the instruction is also considered. This helps the robot focus on the words that matter for the task. Across simulated and real robot experiments, LangForce improves instruction following, especially when visual cues are ambiguous or misleading.