VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos
Abstract
Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions, which is crucial in creating truly original and artistic videos. The challenge lies in finding sufficient training videos with the intended uncommon camera motions. To this end, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. VividCam incorporates multiple disentanglement strategies that isolate camera motion learning from synthetic appearance artifacts, ensuring more robust motion representation and mitigating domain shift. We show that our design synthesizes a wide range of precisely controlled camera motions using surprisingly simple synthetic data. Notably, this synthetic data often consists of basic geometries within a low-poly 3D scene and can be efficiently rendered by engines like Unity. Our video results can be found in https://wuqiuche.github.io/VividCamDemoPage/.
Lay Summary
Videos often use camera motion to create emotion, direct attention, or tell a story. However, current AI video generation models struggle with more unusual or expressive motions, such as searching for an object, switching focus between objects, or creating dramatic camera effects. A major obstacle is that training these models normally requires many realistic videos with the desired camera motions, which are difficult and expensive to collect. We introduce VividCam, a method that teaches video generation models new camera motions using simple synthetic videos rendered in virtual 3D scenes. Although these training videos look unrealistic, VividCam is designed to learn the camera movement while avoiding the visual style of the synthetic data. As a result, the model can generate realistic-looking videos that follow more diverse and precise camera controls. This makes it easier to create expressive videos with customized camera movements, without requiring large collections of real training videos.