Component-Wise Composite Likelihood Distillation for Censored Time-to-Event Data
Abstract
Accurate survival modeling in biomedical studies is often hindered by rare events, limited effective sample sizes, and settings with limited or partially observed information (e.g., covariates of interest that are difficult or expensive to collect, highly structured sampling designs, or nuisance parameters omitted by conditioning). Knowledge distillation can leverage external predictive information without sharing individual-level data, but existing approaches are largely built for fully specified likelihoods or probability-based survival models and do not extend to settings where outcome distributions are only partially specified. To address this challenge, we propose a knowledge distillation framework based on a composite-likelihood Kullback--Leibler divergence that aligns teacher and student models within components. Our key insight is that, although composite likelihoods do not define a global outcome distribution, each likelihood component induces a well-defined probability model on its restricted outcome space, enabling a principled KL divergence. Simulation studies and biomedical case studies show improved discrimination and predictive accuracy in rare-event, heterogeneous settings without requiring access to external individual-level data.
Lay Summary
Many medical studies aim to predict when important health events may occur, such as disease progression, transplant failure, or death. This is difficult when events are rare, when only a small number of patients are available, or when some useful measurements are collected only for selected people because they are costly or hard to obtain. Nested case-control studies are a common example: instead of measuring everyone in a large cohort, researchers measure patients who experience an event and a small set of comparable patients who do not. These designs often rely on composite likelihood methods, which combine many small, local comparisons rather than modeling the full patient population at once. This makes standard knowledge distillation methods hard to use, because the student model does not learn from ordinary patient-level outcome probabilities. We propose a new distillation method that lets a student model learn from an external teacher model within these local comparison groups. This allows useful external knowledge to be transferred without sharing the original external patient-level data. In simulations and transplant-registry studies, our approach improved risk prediction in rare-event and data-limited settings.