From Offline Trajectories to Online Adaptation: A Multimodal JEPA Pretraining Study on Pokemon Red
Stefano Campese ⋅ Alessandro Moschitti
Abstract
We study self-supervised pretraining for reinforcement learning in a long-horizon, partially-observable environment (Pokemon Red). A multimodal JEPA encoder fusing pixels and game RAM is pretrained on offline trajectories and used as a frozen feature extractor for downstream PPO. On 48 held-out starting states never seen during training, frozen JEPA outperforms generic vision pretraining (DINOv3), random initialization, and full-encoder fine-tuning by 2--8$\times$ in cumulative reward. Initializing JEPA from DINOv3 weights yields OOD performance comparable in mean to end-to-end DreamerV3 ($4.48 \pm 0.15$ vs $4.19 \pm 0.80$), with $\sim 5\times$ lower cross-seed standard deviation, indicating substantially higher reproducibility under a simpler two-stage pipeline. Linear probes show fine-tuning catastrophically degrades the pretrained representation: event-progression MSE rises $3.7\times$ while map and party-HP probes lose 5--68\%. A data ablation reveals that data *diversity* matters more than data *expertise*: trajectories collected by a biased-random policy from diverse starting states yield strictly better pretraining data than trajectories from a trained expert policy on the same start distribution.
Video
Chat is not available.
Successful Page Load