Where Does Prediction Error Come From When the Data Is Perfect? A Decomposition of the Model–World Gap in Predictive Uncertainty
Abstract
Most discussions of predictive uncertainty in machine learning focus on data problems, e.g. finite samples, measurement error or distribution shift, as the dominant sources of error. We argue that uncertainty in predictions is structured, even when the analyst has access to large random samples from the target distribution. We refer to this as the model--world gap and, drawing together threads from the statistical, sociological, and ML uncertainty-quantification literatures, develop a layered decomposition that distinguishes five distinct error sources: aleatoric variability, concept-induced inflation, hypothesis class misspecification, asymptotic estimator bias, and finite-sample error. The framework provides a conceptual foundation for thinking about predictive uncertainty independently of data quality, a first step toward more principled diagnosis and mitigation of prediction error in high-stakes applications.