FICReg: Forward-Inverse Consistency Regularization for Latent World Models
Abstract
Joint Embedding Predictive Architecture (JEPA) based world models learn dynamics efficiently in latent space. However, their forward predictors are typically supervised to match target states, without an explicit mechanism to verify that predictions preserve the causal effect of actions. While Inverse Dynamics Models (IDMs) have been used to learn action-relevant representations, they are typically trained jointly with the encoder on ground-truth states, risking co-adaptation and shortcut solutions. We propose the Forward-Inverse Consistency Regularizer (FICReg), which repurposes the IDM as a frozen consistency judge for the forward predictor: the IDM is trained on ground-truth transitions but applied to predicted states with its weights frozen, so that gradients flow exclusively into the predictor. This controlled gradient path improves the dynamics learned by the forward model—pushing it to produce states from which the executed action is recoverable—without distorting the encoder or the IDM itself. Preliminary experiments on four continuous control environments suggest that FICReg benefits tasks where actions play a decisive causal role in state transitions.