Imagined Memorisation: Training-Data Leakage in Model-Based RL World Models
Wang Ngai Ng
Abstract
World models such as DreamerV3 and IRIS are action-conditioned long-horizon video generators: from an encoded context they imagine 15--45 future frames under a user-supplied action sequence, a setting structurally analogous to controllable video generation. We present the first systematic membership-inference audit of these models, adapting three attack families (trajectory reconstruction, dynamics-loss MIA, and adversarial-action divergence) to the action-conditioned generative setting and evaluating across DreamerV3 and IRIS on four Atari environments. On IRIS / Ms.\ Pac-Man, reconstruction-based MIA attains AUC$=0.999$ at $H{=}45$ with Cohen's $d=-4.76$ and TPR$=0.98$ at $1\%$ FPR, exceeding signals typically reported for language and diffusion models; yet on the same checkpoint the standard loss-based MIA flags zero members, and five of eight loss-MIA evaluations score below random. We attribute the disagreement to a collection-policy state-space mismatch that swamps likelihood-based scores while leaving pixel-level signals intact. The implication is that memorisation in pixel-generative video models concentrates in the decoder pathway---a finding that bears directly on how long-horizon video generators should be audited and evaluated.
Chat is not available.
Successful Page Load