A Strictly Proper Scoring Rule and a Calibration Metric for Interval-Censored Data Analysis
Abstract
Lay Summary
Many real-world prediction problems involve event times that can't be observed exactly. For example, in medical studies, an infection might be detected only at periodic visits, so the true onset time is known only to lie between two dates. This is called interval censoring, and it makes it hard to both train and evaluate probabilistic models because the usual likelihood and calibration tools assume exact labels. In this paper, we provide the first strict theoretical guarantee that the standard interval-censored log loss is a strictly proper scoring rule under some common assumptions. In addition, we introduce a new calibration metric that can be computed from censored intervals and detects when predicted probabilities are systematically too optimistic or pessimistic. Experiments show a neural network trained with the proposed loss is competitive with strong statistical baselines.