FAST-AR: Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Abstract
Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their core attention layers become a major bottleneck at inference time: as generation progresses, the KV cache grows, causing both increasing latency and escalating GPU memory, which in turn restricts usable temporal context and harms long-range consistency. In this work, we study redundancy in autoregressive video diffusion and identify three persistent sources: near-duplicate cached keys across frames, slowly evolving (largely semantic) queries/keys that make many attention computations redundant, and cross-attention over long prompts where only a small subset of tokens matters per frame. Building on these observations, we propose a unified, training-free attention framework (FAST-AR) for FAST-AutoRegressive diffusion, consisting of three components: TempCache compresses the KV cache via temporal correspondence to bound cache growth; AnnCA accelerates cross-attention by selecting frame-relevant prompt tokens using fast approximate nearest neighbor (ANN) matching; and AnnSA sparsifies self-attention by restricting each query to semantically matched keys, also using a lightweight ANN. Together, these modules reduce attention, compute, and memory and are compatible with existing autoregressive diffusion backbones and world models. Experiments demonstrate up to x5--x10 end-to-end speedups while preserving near-identical visual quality and, crucially, maintaining stable throughput and nearly constant peak GPU memory usage over long rollouts, where prior methods progressively slow down and suffer from increasing memory usage.
Lay Summary
Generating long videos with AI is becoming increasingly important for applications such as virtual worlds, interactive games, and simulation systems. However, today’s video generation models often become slower and use more GPU memory the longer they run, because they keep storing and reusing more information from previous frames. This makes it difficult to generate long, consistent videos efficiently. In this work, we show that much of this stored and recomputed information is repetitive. Many video frames contain very similar visual content, many internal model calculations change only slowly over time, and long text prompts often contain many words that are irrelevant to a particular frame. Based on these observations, we introduce FAST-AR, a training-free method that makes autoregressive video generation faster and more memory-efficient. FAST-AR compresses repeated information from past frames, focuses text-based attention on the most relevant prompt words, and limits video attention to the most useful past information. Across experiments, FAST-AR speeds up video generation by up to 5–10× while preserving nearly the same visual quality. It also keeps GPU memory and generation speed stable during long videos, making long-form and interactive video generation more practical.