Veda: Scalable Video Diffusion via Distilled Sparse Attention
Abstract
Lay Summary
Video diffusion models have rapidly advanced toward producing visually faithful videos with rich detail, coherent motion, and increasingly realistic scene dynamics. This progress, however, exposes a fundamental scaling challenge: as videos become longer and higher in resolution, generation requires coordinating far more information across space and time, making the process prohibitively slow. Veda addresses this bottleneck by reducing redundant computation while preserving the essential internal structure that the original model relies on. Rather than simply removing computation, Veda learns which interactions are structurally important and keeps those pathways, allowing the model to spend computation where it matters most. This leads to substantial practical acceleration without noticeable degradation in visual quality. For a 720p, 10-second video with 241 frames, Veda reduces generation time from 19.4 minutes to 3.8 minutes, achieving a 5.1× end-to-end speedup. Since the benefit increases as videos become longer and higher-resolution, Veda provides a concrete path toward practical high-resolution, long-duration video generation.