Physics-Guided Motion Loss for Video Generation Model
Abstract
Current video diffusion models generate visually compelling content but often struggle with physical motion, producing subtle artifacts like rubber-sheet deformations and inconsistent object motion. We introduce a frequency-domain physics prior that improves motion plausibility without modifying model architectures. Our method decomposes common motion patterns (translation, rotation, scaling) into lightweight spectral losses. Applied to Open-Sora, MVDIT, and Hunyuan, our approach improves both motion accuracy and action recognition by ∼11\% on average on OpenVID-1M (relative), while maintaining visual quality. Additional results on Wan 2.1-14B show consistent gains on video-quality and physics-oriented metrics. User studies show 74-83\% preference for our physics-enhanced videos. It also reduces warping error by 22-37\% (depending on the backbone) and improves temporal consistency scores. These results indicate that simple, global spectral cues are an effective drop-in regularizer for physically plausible motion in video diffusion.
Lay Summary
AI video generation systems can create visually impressive videos, but their motion is often not physically realistic. Objects may stretch, flicker, rotate inconsistently, or change size in unnatural ways. This paper presents a simple training method that helps video models produce smoother and more physically plausible motion. It encourages generated videos to follow basic motion patterns, such as steady movement, consistent rotation, and coherent changes in scale. The method works with existing video generation systems and improves motion quality while preserving visual appearance.