Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
Abstract
Lay Summary
Video generation models can now create high-quality videos, but they are still slow and expensive to run. This paper introduces Light Forcing, a method that makes autoregressive video generation faster by reducing unnecessary computation while keeping important past information. Unlike previous sparse attention methods, Light Forcing is designed for the step-by-step generation process of autoregressive models. It decides how much computation each video chunk needs and selects useful past frames and regions at different levels of detail. Experiments show that Light Forcing maintains strong video quality while achieving 1.2–1.3× end-to-end speedup, and can further reach 2.0–3.0× speedup when combined with other efficiency techniques, making high-quality video generation more practical.