What Matters for Music-Centered Recognition in Audio-Language Models?
Abstract
Audio-language models (AudioLMs) offer a flexible natural-language interface for open-ended audio understanding, yet their efficacy on structured, closed-set audio recognition remains poorly understood. This paper systematically investigates this question across four diverse recognition tasks, with a primary focus on music-centered domains. Internal probing shows that task-relevant information often remains accessible after the audio encoder, while controlled comparisons show that LLM-based label generation changes recognition behavior in a task-dependent way. Specifically, the language interface consistently benefits environmental sound recognition but degrades music genre tagging compared to direct encoder classification. Systematic evaluations of model design choices further establish small-learning-rate encoder tuning as the most robust adaptation strategy, while attention-based multi-encoder fusion gives the strongest in-domain joint model. These results suggest that closed-set AudioLM recognition is limited less by a simple loss of audio information than by how audio representations are adapted, combined, and decoded into fixed labels.