Toward Continually Improving Long-Horizon Agents
Abstract
Long-horizon agents need to learn from evolving data, remember useful experience, and use that experience to improve future behavior. This talk presents a research trajectory toward such continually improving agents through three connected stages. I will first discuss how agents can selectively expand their capabilities as new multimodal instruction data becomes available, avoiding redundant learning while acquiring a broader range of skills. I will then move to the problem of memory, where long-form visual experiences require agents to preserve both abstract knowledge and fine-grained evidence across extended temporal contexts. Finally, I will discuss how agents can go beyond passive understanding by using their own failures to generate targeted environments for embodied skill acquisition. These directions outline a path toward agents that can select what to learn, organize what they remember, and improve their behavior over time.