Physically Grounded Video-to-Audio Generation
Abstract
Video-to-audio (V2A) models can now synthesize perceptually plausible and temporally aligned sounds, but they often remain weakly grounded in the physical quantities that shape real acoustic events, such as object mass and motion. We introduce PAVAS, a physics-aware V2A framework that estimates object-level mass and velocity from video and injects these cues into a latent diffusion backbone through a lightweight Physics-Driven Audio Adapter. To evaluate this behavior, we curate VGG-Impact, a subset of impact-centric VGGSound clips, and propose the Audio-Physics Correlation Coefficient (APCC), which measures whether generated onset strengths vary consistently with estimated kinetic-energy changes. Across VGGSound and VGG-Impact, PAVAS improves both standard metrics and physics consistency compared with prior strong V2A models.