What Matters in Clean-Context Autoregressive Video Diffusion
Abstract
Long-horizon autoregressive video diffusion drifts. Recent methods close this gap by training with a clean past prefix and masking its loss, the video analogue of standard autoregressive language-model training in which the model is supervised only on the conditional it will be queried for at inference. We provide a compute-matched controlled study on the public Diffusion Forcing codebase, isolating the two training-side axes (the clean prefix and the masked prefix loss) under two complementary distributional metrics on rollouts: FVD and JEDi. In this controlled DMLab setting the clean prefix is the dominant axis: it substantially reduces both endpoint and segment-wise drift under both metrics, and the additional benefit of loss-masking shows up under FVD but does not replicate under JEDi, under which the clean-only variant performs the best at every segment of the rollout. Long-horizon video evaluation should therefore report a temporally decomposed metric and at least one alternative-backbone replication, since endpoint FVD alone obscures both effects.