ListenCare: Encounter-Grounded Audio Question Answering for Long-Form Clinical Conversation Speech
Abstract
Clinical conversation speech is a rich source of evidence for ambient clinical documentation, patient-facing summaries, and care coordination. Yet existing audio-language benchmarks rarely test whether models can answer clinically relevant questions grounded in long-form doctor-patient encounter audio. We introduce ListenCare, an encounter-grounded audio question answering benchmark with 4,085 four-option MCQA instances from 457 synthetic and mock clinical encounters. The benchmark is organized by a six-capability taxonomy spanning evidence recovery, evidence status, attribution grounding, temporal/discourse tracking, interaction-level reasoning, and audio-specific grounding. We evaluate five open-weight large audio-language models (7B-30B parameters) under full-audio QA, ASR-transcript Cascade, and oracle evidence-localized audio settings. We find two main results: less capable LALMs still suffer from audio-text modality gaps and from retrieving evidence in long audio, whereas stronger models reduce these bottlenecks and achieve higher accuracy. Even the strongest models, however, remain challenged on interaction-level reasoning and audio-specific grounding, pointing to interactional and acoustic understanding beyond retrieving spoken content as the next bottleneck.