SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
Abstract
We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.
Lay Summary
Adding sound to silent video is difficult because the audio must not only sound realistic, but also happen at exactly the right moment. We built SALSA-V, a deep learning model that can create synchronized audio for a given video, such as footsteps, impacts, voices, or environmental sounds, while keeping those sounds aligned with the visible events. Unlike many previous systems that work best on short clips, SALSA-V can extend audio over longer videos by using earlier generated sound as context. It can also use a short example sound to match a desired style or recording quality, giving users more control over the result. The model is designed to produce high-quality audio in only a few generation steps, making fast feedback possible. This could make sound design easier for films, games, user-generated videos, and accessibility tools.