Hidden Failure Modes in Latent World-Model Planning from Offline Data
Kanpat Vesessook ⋅ Kevin Yang
Abstract
Latent world models offer an appealing route from offline datasets to online decision-making: learn compact predictive dynamics from logged trajectories, then adapt online by planning in latent space.
We revisit LeWorldModel-style joint-embedding predictive architectures under closed-loop receding-horizon MPC, where a planner optimizes horizon $H$ but executes only a prefix $K < H$ before replanning.
Across standard tasks, we find that terminal-at-$H$ scoring can misdiagnose model quality because the score is evaluated at a time index the controller never directly executes: prefix-terminal scoring repairs PushT and TwoRoom, while running costs substantially repair the broader standard-task set.
On deceptive navigation, however, aligned scalar latent costs remain insufficient: the best scalar/value-bootstrapped latent objective reaches $26.5\%$, while a waypoint/local-actuator interface reaches $78.8\%$--$92.7\%$ on held-out TwoRoom Far Door, with nearest retrieval remaining a strong control.
The main lesson is that offline-to-online world-model evaluations can fail at the planning interface even when the learned representation is usable.
These results suggest that offline-to-online latent planning benchmarks should report not only learned model quality, but also replanning interval, scoring time index, and controllability interface.
Chat is not available.
Successful Page Load