Toward Compositional Latent Action Interfaces for Generalizable Agents
Abstract
Latent Action Models (LAMs) learn action proxies from observation transitions, but they face a fundamental ambiguity in multi-object or distractor-rich scenes: without supervision, the model cannot determine which changes are caused by the controlled agent. When multiple transition sources and environment-specific factors are compressed into a single monolithic latent action, the resulting representation can become sensitive to distractors and may generalize poorly under distribution shift. We introduce Observed Transition Factorization (OTF), a framework that represents each transition as a sparse composition of reusable observed-transition primitives, yielding a compositional latent action interface. Building on this representation, we propose OTF-LAM, a latent action model that constructs action-like latents from factorized transition primitives within the standard inverse-forward dynamics framework. Empirically, we show that the learned transition primitives transfer zero-shot across both controlled setting and Distracting Control Suite environments, and can be used for downstream policy learning. Our code will be at anonymous.