Inference-Time On-Manifold Steering for Autoregressive Long Video Generation
Abstract
Autoregressive video diffusion models have shown remarkable success in real-time video generation. However, generating long video sequences with these models remains challenging due to the progressive accumulation of errors and temporal drift. We attribute this degradation to the sequential generation process, which gradually pushes latent representations off the desired manifold. To address this, we introduce On-Manifold Steering with Time Reward, a novel framework designed for long video generation. By training a lightweight time predictor to classify the timesteps of noisy latents, we derive a time reward gradient that continuously steers the diffusion sampling trajectory. This guidance actively pulls the temporal latent states back onto the manifold at each step, effectively correcting inference misalignments on the fly. Consequently, our approach mitigates feature drift and preserves structural integrity across long video sequences. Extensive experiments demonstrate that our framework seamlessly integrates with existing baselines such as Self Forcing and Deep Forcing, consistently improving temporal coherence and visual quality in 30- and 60-second video generation tasks.