DialBGM: A Benchmark for Background Music Recommendation from Everyday Multi-Turn Dialogues
Abstract
Selecting suitable background music (BGM) for natural human conversation is a common yet underexplored production step in media and interactive systems. We introduce dialogue-conditioned BGM recommendation, a task that requires select- ing non-intrusive, suitable music for a multi-turn conversation that typically contains no music descriptors. To study this problem, we present DialBGM, a benchmark of 1,200 open-domain daily dialogues, each paired with four candidate music clips and annotated with human preference rankings. We evaluate audio–language models, multimodal LLMs, and embedding-based retrieval models, and find that no model exceeds 35% Hit@1, indicating a substantial gap relative to human judgments. DialBGM provides a standardized testbed for developing affective-reasoning methods for BGM selection.