Hard-First: Entropy-Guided Curriculum Distillation Balances Transfer and Preservation in Biomedical Vision-Language Models
Junseob Kim ⋅ Sunil Hwang ⋅ Mehak Arora ⋅ Rishikesan Kamaleswaran
Abstract
Knowledge distillation (KD) offers a natural path to compress large biomedical vision-language models into smaller students that meet the compute, latency, and data-locality constraints of clinical deployment. However, KD suffers from catastrophic out-of-distribution (OOD) forgetting, where fine-tuning on a narrow mix degrades generalization to unseen modalities and task structures. We introduce Hard-First, an entropy-guided curriculum that orders samples from highest to lowest teacher predictive entropy, prioritizing hard examples to better preserve decision-boundary information. To evaluate both transfer and retention, we design a two-axis OOD framework that separates visual-modality shifts (unseen imaging modalities; MedBookVQA) from structural shifts (multi-image reasoning; MedFrameQA). Across 10 configurations on five benchmarks with Qwen2.5-VL-72B$\to$3B, Hard-First achieves the largest transfer gain (PathVQA $+8.6$ pp over supervised fine-tuning) while minimizing OOD degradation (MedBookVQA $-2.7$ pp, MedFrameQA $-4.5$ pp). Cross-family replication (InternVL3-38B$\to$2B) further establishes Hard-First as the top OOD-preserving ordering.
Chat is not available.
Successful Page Load