Auditing Clinical Concept Fragmentation in Sparse Medical Vision–Language Representations
Abstract
Trustworthy clinical AI requires that model evidence be inspectable at the level of clinically meaningful concepts, not only individual sparse features. Sparse dictionary models expose internal activations, but in medical vision–language models they can split one coherent finding across many atoms, creating a failure mode for clinical auditing. We study this failure mode as concept fragmentation and measure it with the Concept Fragmentation Score (CFS), the effective number of sparse features used by each supported clinical concept. We introduce HARP, a hierarchy-aligned, report-guided Poincaré objective that aligns sparse image codes with report-derived UMLS targets at the study level. On held-out MIMIC-CXR, HARP reduces CFS from 76.9 to 22.3 relative to a per-feature Euclidean ontology baseline while preserving reconstruction. The same frozen MIMIC-trained dictionary reduces CFS on CheXpert, NIH ChestX-ray14, and OpenI; linear probes improve on NIH and OpenI and remain comparable on CheXpert. A per-feature Poincaré prototype diagnostic collapses and worsens CFS, showing that ontology metric and supervision granularity must be evaluated together.