Cinematic Source Separation with Dialogue-Driven Sidechain Ducking
Abstract
Cinematic audio source separation, the task of extracting dialogue, music, and effects stems from film soundtracks, is a prerequisite for speech-centric post-production workflows including dubbing, accessibility captioning, and broadcast ASR, yet is limited by the lack of realistic training and evaluation data. Existing synthetic datasets sum stems linearly without production artifacts such as sidechain ducking, where music and effects are attenuated in the presence of speech to protect dialogue intelligibility. Real mixes apply this ducking on the mix bus, so the original stems no longer sum to the output mix. We present CineAudioGen, an agentic pipeline that generates four-stem cinematic training data using a Director-Critic architecture. We also release CineAudioDB, a dataset of real cinematic audio with ground-truth stems from independent filmmakers. Fine-tuning Bandit v2, MRX, and HTDemucs on our data yields considerable improvements on cinematic content compared to baselines.