LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
Abstract
Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propose \textbf{LUVE}, a \textbf{L}atent-cascaded \textbf{U}HR \textbf{V}ideo generation framework built upon dual frequency \textbf{E}xperts. LUVE employs a three-stage architecture comprising low-resolution motion generation for motion-consistent latent synthesis, video latent upsampling that performs resolution upsampling directly in the latent space to mitigate memory and computational overhead, and high-resolution content refinement that integrates low-frequency and high-frequency experts to jointly enhance semantic coherence and fine-grained detail generation. Extensive experiments demonstrate that our LUVE achieves superior photorealism and content fidelity in UHR video generation, and comprehensive ablation studies further validate the effectiveness of each component.
Lay Summary
Recent AI tools have gotten very good at generating videos, but creating ultra-high-resolution (UHR) videos is still incredibly difficult. This is because the AI has to simultaneously figure out realistic movement, keep the overall scene making sense, and draw incredibly fine details. To solve this, we created a new AI system called LUVE. Instead of trying to do everything at once, LUVE breaks the video creation process into three manageable steps. First, it generates a small, low-resolution version of the video just to get the movement and flow right. Second, it efficiently enlarges this video in a way that doesn't overwhelm the computer's memory. Finally, it acts like a team of two specialized artists: one ensures the overall scene makes sense, while the other draws in all the sharp, highly realistic details. Our tests show that this step-by-step approach allows LUVE to generate stunning, highly realistic ultra-high-resolution videos.