MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
Abstract
Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we present MedMosaic, a medical audio question–answering dataset designed to benchmark language and audio reasoning models under realistic clinical constraints. MedMosaic features a diverse range of medical audio types, including condition-related physiological sounds, carefully constructed synthetic voices to mimic speech with artifacts as well as real short and long length clinical conversations to model varying context lengths. The dataset also features a total of 46,701 question-answer pairs, spanning categories such as multiple-choice, sequential multi-turn, and open-ended question–answers, enabling systematic evaluation of multi-hop reasoning and answer generation capabilities. Benchmarking 13 audio and multimodal reasoning models reveals that reasoning remains challenging for all evaluated systems, with substantial performance variation across question types. In particular, even state-of-the-art model like Gemini-2.5-pro can only achieve 68.1\% accuracy approximately. These findings underscore persistent limitations in medical reasoning and highlight the need for more robust, domain-specific multimodal reasoning models. A sample of benchmark data is available here:https://shorturl.at/Lyp33
Lay Summary
Doctors rely heavily on what they hear: a heart murmur through a stethoscope, the character of a cough, or subtle hesitations in a patient's voice during a long conversation. These acoustic signals carry vital diagnostic clues, yet we have no good way to measure whether AI systems can actually understand and reason over medical audio. The core barrier is data scarcity, as privacy regulations and the high cost of expert annotation make it extremely difficult to build large-scale evaluation resources in this domain. We created MedMosaic, a benchmark of over 46,000 question–answer pairs spanning heart sounds, lung sounds, coughs, clinical conversations of varying lengths, and recordings that blend speech with physiological sounds. The benchmark draws on diverse publicly available medical audio datasets and is supplemented by a synthetic audio generation pipeline that produces realistic clinical recordings enriched with physiological cues such as coughs, wheezes, and vocal distress markers. Healthcare professionals validated the resulting questions with a 72% acceptance rate. When we tested 13 state-of-the-art AI models, the best one answered correctly only 68% of the time, and most performed far worse, exposing fundamental weaknesses in how current AI handles medical audio reasoning. MedMosaic gives the research community a rigorous, publicly available tool to drive progress toward AI systems that can one day support clinicians in making faster, more accurate diagnoses from the sounds at the bedside.