A Machine-Learned Comorbidity Index
Abstract
Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outcomes, and (ii) their linear, rule-based structure cannot capture nonlinear, outcome-specific risk relationships. We propose a Machine-Learned Comorbidity Index (MLCI) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes. MLCI captures nonlinear risk–outcome dependence and is supported by a theory that characterizes when a unified, informative admission-level ordering can be achieved across outcomes. Empirical results on multiple benchmark electronic health record (EHR) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics.
Lay Summary
We introduce Machine-Learned Comorbidity Index (MLCI), a data-driven way to summarize how severe a patient’s condition is during a hospital admission from their diagnosis codes. Unlike traditional scores, MLCI is not limited to fixed rules or primarily designed on mortality risk. It learns one simple severity score that is useful across several outcomes, including death, hospital length of stay, and ICU transfer. Tests on two large electronic health record datasets show that MLCI captures patient outcomes better than existing comorbidity scores and several machine-learning baselines.