Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering
Xupeng Chen ⋅ Binbin Shi ⋅ Chenqian Le ⋅ Jiaqi Zhang ⋅ Kewen Wang ⋅ Ran Gong ⋅ Jinhan Zhang ⋅ Chihang Wang
Abstract
Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich visual content (tables, figures, structured layouts) of original document pages. We propose \MedVRAG{}, an iterative multimodal RAG framework that retrieves and reasons over PMC document page images instead of OCR'd text. The system pairs ColQwen2.5 patch-level page embeddings with a sharded MapReduce LLM filter, scaling to $\sim$350K pages while keeping Stage-1 retrieval under 30\,ms via an offline coarse-to-fine index ($C{=}8$ centroids per page, ANN over centroids, exact two-way scoring on the top-$R$ shortlist). A vision-language model (VLM) then iteratively refines its query and accumulates evidence in a memory bank across $\le$3 reasoning rounds, with a single iteration costing $\sim$15.9\,s and the full three-round pipeline $\sim$47.8\,s on 4$\times$A100. Across four medical QA benchmarks (MedQA, MedMCQA, PubMedQA, MMLU-Med), \MedVRAG{} reaches 78.6\% average accuracy. Under controlled comparison with the same Qwen2.5-VL-32B backbone, retrieval contributes a $+5.8$-point gain over the no-retrieval baseline; we also note a $+1.8$-point edge over MedRAG\,+\,GPT-4 (76.8\%), with the caveat that this is a cross-paper rather than head-to-head comparison. Ablations isolate $+1.0$ from page-image vs.\ text-chunk retrieval, $+1.5$ from iteration, and $+1.0$ from the memory bank.
Chat is not available.
Successful Page Load