Efficient Multimodal Clinical Question Answering for Pulmonary Embolism Risk Assessment
Abstract
Pulmonary embolism (PE) is a high risk cardiopulmonary condition whose management requires both timely diagnosis and reliable assessment of future clinical risk. Because PE care routinely combines computed tomography pulmonary angiography (CTPA), radiology interpretation, and longitudinal electronic health record (EHR) evidence, it provides a clinically meaningful setting for evaluating compact multimodal language models. In this work, we build a benchmark using efficient vision language models (VLMs) on INSPECT, a multimodal PE dataset containing 23,248 CTPA studies from 19,402 patients. We formulate eight diagnostic and prognostic tasks as structured clinical question answering problems and evaluate on typical efficient VLMs under CT only, EHR only, and CT plus EHR settings with zero and few shot prompting. Results show that Gemma series models perform more strongly when EHR evidence is available, especially under CT plus EHR input. Task level analysis further shows that PE diagnosis is much easier than longitudinal prognosis, particularly readmission prediction. These observations suggest that compact multimodal models have the great potential in early stage PE risk detection and explanation.