Beyond Accuracy: Epistemic Justification in Trustworthy Machine Learning
Abstract
Machine learning systems are routinely certified as "trustworthy" on the basis of predictive accuracy, yet epistemology draws a sharp distinction between a belief being true and a belief being justified. We argue that, under standard epistemological frameworks, a model can be predictively successful without satisfying the stronger justificatory criteria that epistemically legitimate prediction requires. We formalize one tractable dimension of this gap via the Justification Deficit (JD), a diagnostic quantity measuring the degree to which a model's confident predictions depend on causally irrelevant features. Our contribution is not a new causal training objective, but a philosophically interpretable evaluative lens on when predictive confidence is epistemically suspect. Experiments on two benchmarks with known causal structure show that high benchmark accuracy does not guarantee low JD, and that standard training routinely sacrifices epistemic legitimacy for predictive performance.