When Can We Trust Survival Model Evaluation ?
Abstract
Evaluating survival models under censoring is inherently challenging, yet standard evaluation practices are often applied without explicitly assessing how censoring distorts metric reliability. Performing a large experimental study, we analyze and quantify how survival evaluation metrics are affected in fundamentally different ways by the censoring rate and the censoring mechanism. Using a controlled semi-synthetic framework, we vary both the censoring mechanism (administrative, independent, covariate-dependent) and the censoring rate, and compare standard evaluations based on censored data with oracle evaluations using fully observed event times. This controlled setting enables us to quantify distortions along two complementary axes: numerical bias and preservation of model ranking. Across datasets and metric families, we find that censoring induces systematic, mechanism-dependent distortions. Moderate numerical bias, if not properly addressed, can lead to unreliable model comparison as censoring increases. These findings reveal fundamental limitations of common benchmarking practices and call for more careful interpretation of survival evaluation under realistic censoring.
Lay Summary
Many real-world prediction problems involve estimating when an event may happen, such as when a patient may relapse, when a customer may leave a service, or when a machine component may fail. These are called survival prediction problems. They are difficult because, very often, we do not observe the event for everyone: for some individuals, we only know that the event had not happened yet. This missing information, called censoring, can make model evaluation misleading. In this paper, we study how much censoring can distort the way survival prediction models are evaluated. We create controlled versions of real datasets where we know the true event times, then compare normal evaluations based on censored data with “oracle” evaluations that use the hidden true times. This allows us to measure when common evaluation scores become biased and when they can incorrectly suggest that one model is better than another. We find that heavier censoring often makes evaluation less reliable, especially when censoring is related to the characteristics of individuals. Our results show that researchers should be careful when declaring a “best” survival model and should report censoring conditions, use statistical comparisons, and choose metrics suited to the problem.