VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World Models
Abstract
Joint Embedding Predictive Architectures (JEPAs) avoid pixel reconstruction by predicting latent representations, but standard formulations remain deterministic and provide limited uncertainty estimates for planning and control. We introduce \emph{Variational JEPA (VJEPA)}, a probabilistic extension that learns predictive distributions over future latent states using a latent-space variational objective, without autoregressive observation likelihoods. We show that VJEPA links JEPA-style self-supervised learning to predictive state representations and Bayesian filtering, and that its latent variables can serve as sufficient information states for control when they preserve task-relevant predictive information. We also propose \emph{Bayesian JEPA (BJEPA)}, which combines a learned dynamics expert with modular prior experts through a Product of Experts, enabling constraint-aware prediction and zero-shot prior swapping. Experiments on Noisy-TV systems, nonlinear and image-based benchmarks, STL-10 with a ViT encoder, and DMC Cheetah-run show that predictive JEPA-family objectives are more robust to high-variance nuisance distractors than reconstruction-based world-model baselines. These results position probabilistic latent prediction as a principled framework for robust, uncertainty-aware, reconstruction-free world models.
Lay Summary
AI systems that plan or make decisions need to predict what may happen next. Many current world models learn by reconstructing detailed observations, but this can waste capacity on irrelevant noise or distractions. We propose VJEPA, a model that instead predicts useful hidden descriptions of future states and estimates its uncertainty. We also introduce BJEPA, a modular extension that can combine learned dynamics with goals or constraints. Across noisy prediction, image-based, and control experiments, these predictive models are more robust to distracting high-variance noise than reconstruction-based baselines. This suggests that probabilistically predicting useful future representations, rather than rebuilding raw observations, is a promising route to robust and uncertainty-aware world models.