Juhan Nam: LLM4FM: Empowering LLMs to Generate Yamaha DX7 Pat
Abstract
Abstract: Programming FM synthesizers is notoriously difficult due to the highly non-linear relationship between parameters and sound. In this talk, I present LLM4FM, a framework that enables Large Language Models (LLMs) to generate Yamaha DX7 patches from text descriptions or audio examples. To support this task, we introduce DX7Caps, the first dataset pairing DX7 patches with natural-language captions, and propose Operator-Isolated Audio Grounding for CoT Distillation. We further extend the framework to sound matching by generating DX7 patches directly from audio. Results from objective evaluations, listening tests, and LLM-based assessments demonstrate the potential of LLMs as practical assistants for FM sound design.
Bio: Juhan Nam is a Professor at the Graduate School of Culture Technology at KAIST and directs the Music and Audio Computing Lab. His research focuses on music information retrieval, audio and music signal processing, generative audio, and human–AI musical interaction.