A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Abstract
Lay Summary
Imagine listening to a recording of a busy street and being able to remove everything except the dog barking, or keep only the rain. AI systems that pull individual sounds out of such mixtures are powerful, but they often leave traces of background noise behind. The reason hides in their training data. Most clips labeled "rain" also contain wind, traffic, or chatter, so the AI quietly learns those extras as part of what rain sounds like. The usual fix has been to throw vastly more data at the problem, training on millions of hours of audio with the computing budget to match. We took the opposite approach. Using modern AI models as automatic auditors, we sifted through twelve public audio collections and kept only short clips containing a single, clearly identifiable sound. We then recombined these clean clips into realistic mixtures, letting "typing" co-occur with "air conditioning" while rejecting impossible pairings such as "whales" with "city traffic." The resulting dataset, named Hive, is roughly 500 times smaller than the largest competitor, yet separation models trained on it match or beat that competitor and generalize well to unfamiliar recordings. Cleaner ingredients, it turns out, can replace bigger ovens.