Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
Woojung Han ⋅ Seil Kang ⋅ Youngjun Jun ⋅ Min-Hung Chen ⋅ Fu-En Yang ⋅ Seong Jae Hwang
Abstract
Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws. We reveal a surprising finding: a 2-step generation often exhibits better physical consistency than a 50-step output from the same model. Through spectral analysis, we trace this to phase erosion during denoising; the phase degrades significantly (dropping by $\approx 18\%$ from step 2 to step 50), whereas the magnitude remains relatively stable. Building on this insight, we propose PhaseLock, a training-free framework that locks the valid motion priors into the denoising trajectory found in few-step inference. Rather than requiring full-step inference to establish physics, PhaseLock extracts a motion prior from just 2 steps and enforces it onto high-fidelity generation via Latent Delta Guidance. Our approach effectively prevents phase degradation, achieving both high visual fidelity and consistently improved physical consistency scores. Extensive experiments demonstrate an average improvement of 5.1 points across diverse models with negligible overhead ($1.06\times$ time, $1.02\times$ memory), eliminating the need for expensive external guidance methods ($\sim5\times$ time).
Chat is not available.
Successful Page Load