FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation
Abstract
Vision–Language–Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. We conduct a systematic stress test of state-of-the-art VLA models and show that performance degrades sharply as demonstrations are reduced, revealing a key weakness of existing adaptation strategies. To address this, we introduce FOCA, a future-oriented conditioning framework for data-efficient VLA adaptation. FOCA combines explicit prediction of task-grounded future interaction embeddings with implicit alignment to future goal observations, enabling long-horizon reasoning in latent space without pixel-level prediction. This formulation naturally supports action-free co-training with synthetic videos from video world models and can be interpreted as learning a future-conditioned value-like representation. Extensive experiments demonstrate FOCA achieves 95.7\% success with 20 demonstrations on LIBERO, improves 7–12\% on RoboCasa, and delivers up to 26\% absolute gains on real robots, establishing a new state of the art in few-shot VLA adaptation.
Lay Summary
Vision–Language–Action (VLA) models allow robots to learn tasks from large amounts of visual and language data, but they struggle when only a few examples are available for a new task. In this work, we show that current methods degrade sharply in such low-data settings, highlighting a key limitation in how they adapt to new situations. To address this, we propose FOCA, a new approach that helps robots better generalize from limited experience by encouraging them to reason about future outcomes while acting. Instead of relying only on past observations, FOCA learns to anticipate future states and align its decisions with likely goals, enabling better long-horizon planning even with little data. Experiments in both simulated and real-world robotic environments show that FOCA significantly improves performance in few-shot settings and achieves state-of-the-art results in data-efficient robotic learning.