From Offline Trajectories to Online Adaptation: A Multimodal JEPA Pretraining Study on Pokemon Red
Stefano Campese ⋅ Alessandro Moschitti
Abstract
We study self-supervised pretraining for reinforcement learning in a
long-horizon, partially-observable environment (Pokemon Red).
A multimodal JEPA encoder fusing pixels and game RAM is pretrained on
offline trajectories and used as a frozen feature extractor for downstream PPO.
On 48 held-out starting states never seen during training, frozen JEPA
outperforms generic vision pretraining (DINOv3), random initialization,
and full-encoder fine-tuning by 2--8$\times$ in cumulative reward.
Initializing JEPA from DINOv3 weights yields OOD performance
comparable in mean to end-to-end DreamerV3
($4.48 \pm 0.15$ vs $4.19 \pm 0.80$), with $\sim 5\times$ lower
cross-seed standard deviation, indicating substantially higher
reproducibility under a simpler two-stage pipeline.
Linear probes show fine-tuning catastrophically degrades the pretrained
representation: event-progression MSE rises $3.7\times$ while map and
party-HP probes lose 5--68\%.
A data ablation reveals that data *diversity* matters more than data
*expertise*: trajectories collected by a biased-random policy from
diverse starting states yield strictly better pretraining data than
trajectories from a trained expert policy on the same start distribution.
Chat is not available.
Successful Page Load