MoVie: Multimodal Video Compression with Text Guidance
Abstract
Lay Summary
Videos take up a large amount of storage and bandwidth, especially when people watch or share them online. Video compression aims to reduce file size while keeping the video visually clear. However, many existing methods mainly focus on pixel-level changes and motion, and may lose important visual details when the bitrate is very low. This paper introduces MoVie, a new video compression method that uses text descriptions to help preserve the most meaningful parts of a video. For example, if a video contains people, animals, or objects, a short description can guide the model to better understand what content should be kept visually clear. MoVie combines this text guidance with efficient video modeling, so it can maintain better visual quality without greatly increasing computation. Our experiments show that MoVie produces videos that look better at low bitrates than several strong existing compression methods. It also uses less computation than many learned video codecs. In addition to automatic evaluation, a human study shows that viewers more often prefer videos compressed by MoVie over those produced by competing methods. These results suggest that using language as guidance can make video compression more perceptually effective and efficient.