Minje Kim: Neural Speech and Audio Coding: Efficient Representations, Emerging Capabilities, and Open Challenges
Abstract
Abstract This talk will provide an overview of the rapidly evolving landscape of neural speech and audio coding (NSAC). Recent progress has shown that data-driven coding systems, when paired with appropriate learning objectives and model architectures, can achieve substantial gains in coding efficiency over conventional approaches. Beyond improved representational efficiency, NSAC also opens the door to new capabilities that are difficult to realize with traditional codecs, including personalized speech coding, task-specific representations for audio coding for machines (ACoM), and cascaded residual learning frameworks that make neural codecs more flexible and expressive. The talk will also discuss key challenges that are specific to NSAC, including the computational cost of neural encoders and decoders, the risk of hallucination and other generative artifacts, and the difficulty of balancing perceptual quality with faithful signal reconstruction. Finally, I will highlight recent efforts to address these issues through more efficient model architectures, improved training strategies, and semantic loss functions that better align codec behavior with human perception and downstream machine-listening tasks.
Bio Minje Kim is an Associate Professor in the Siebel School of Computing and Data Science at UIUC. Before that, he was an Associate Professor at Indiana University and an Amazon Scholar. He earned his Ph.D. in Computer Science at UIUC after working as a researcher at ETRI, a national lab in Korea. During his career as a researcher, he has focused on developing machine learning models for audio signal processing applications. He is a recipient of various awards, including the NSF Career Award, Facilitating Learning Excellence Award by Illinois Student Council, IU Trustees Teaching Award, IEEE SPS Best Paper Award, Google and Starkey’s grants for outstanding student papers in ICASSP 2013 and 2014, respectively. He is currently serving as the Chair of the IEEE SPS Audio and Acoustic Signal Processing Technical Committee, and as Senior Area Editor for IEEE TASPro and SPL. He was the General Co-Chair of IEEE WASPAA 2023, and has co-organized various NSAC-related academic events, including the ICASSP 2023 Special Session, the IEEE JSTSP Special Issue, the Low-Resource Audio Coding (LRAC) Workshop at ICASSP 2026, and the NSAC tutorial at Interspeech 2024.