Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking
Abstract
We present VTMR, a two-stage framework for video-to-music recommendation designed to address two fundamental limitations of existing methods: reliance on unimodal video representations and the loss of temporal context from average-pooled global embedding search. To overcome these challenges, VTMR first encodes comprehensive multimodal video and music signals into a tightly aligned joint audio-visual-text representation space. We then deploy a retrieve-and-rerank pipeline: Stage~1 retrieves semantically compatible candidates at scale via global embedding search, and Stage~2 reranks them by attending over the full temporal sequences of both video and music. Evaluated on the video-to-music recommendation task, the multimodal retrieval stage improves R@10 from 14.2 to 15.9 and Median Rank from 75 to 58 over the strongest baseline; the temporal reranker further boosts R@10 to 18.3 and Median Rank to 46, demonstrating complementary gains from richer query encoding and temporal alignment. A human preference study confirms that VTMR is on par with a commercial baseline in overall preference, while outperforming a generative baseline in music quality.