Decouple and Cache: KV Cache Construction for Streaming Video Understanding
Abstract
Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and evicting old key-value(KV) caches is required for unbounded streams. Secondly, due to the high cost of collecting and training on unbounded streams, models must learn from short sequences while generalizing to long streams. Existing streaming VideoVLLMs fail to scale to unbounded video streams or focus on cache reuse strategies, leaving the impact of cache construction underexplored. In this paper, we propose Decoupled Streaming Cache(DSCache), a training-free cache construction mechanism that adapts pretrained offline models to streaming settings. DSCache maintains a cumulative past KV cache while constructing a separate instant cache on-demand, decoupled from past caches to preserve the informativeness of recent inputs. To enable position extrapolation beyond the training length, DSCache further incorporates a position-agnostic encoding strategy, ensuring KV caches to support unseen positions and preventing position overflow. Experiments on Streaming Video QA benchmarks demonstrate DSCache's state-of-the-art performance, with an average 2.5% accuracy gains over prior methods.
Lay Summary
Modern video AI systems are typically designed for offline use: they process an entire video before answering questions. However, real-world applications like AR assistants, robotics, or surveillance require models to understand continuous live video streams in real time. Major challenges include efficiently maintaining memory over long streams without repeatedly recomputing past information, and scaling models trained on short video clips to unbounded streaming videos. Existing streaming methods mainly focus on compressing or reusing memory caches, but often overlook how these caches are constructed over time and how to support unbounded video streams. This paper introduces DSCache, a training-free streaming framework that improves the cache construction process and enables models trained on short video clips to scale to effectively infinite video streams during inference. Together, these changes improve both efficiency and streaming video understanding performance without retraining the underlying model.