CE$^4$L: Continual Ego, Exo, and Ego-Exo Learning
Abstract
Lay Summary
AI agents such as robots often learn from videos while their tasks and camera viewpoints change over time. For example, the same activity can look very different from a first-person camera and from an external camera, and a model that remembers old tasks may still fail to transfer knowledge across these views. We introduce CE4L, a benchmark that tests continual learning in multi-view video settings, including skill assessment, action segmentation, cross-view matching, and action anticipation and planning. This benchmark helps researchers measure not only whether a model forgets past tasks, but also whether it preserves knowledge that transfers between egocentric and exocentric views. We also propose VISTA, a memory-efficient method that stores task-specific adaptations and automatically selects the most relevant ones using statistics of video representations. Across CE4L, VISTA performs strongly compared with representative continual learning methods. We hope CE4L will help the community build video learning systems that remain stable and transferable in realistic embodied environments.