Counterfactuals Without Worlds: When ML Counterfactual Explanations Are Ill-Posed
Abstract
Counterfactual explanations are often presented as a natural bridge between machine learning and human understanding: if a system rejects an application, it can say which small changes would have produced acceptance. We argue that many such explanations are not false so much as under-specified. A feature vector that flips a classifier answers a syntactic question about a model; when advertised as explanation, recourse, or contestation, it must also answer a modal, causal, and practical question about what could have been different, why it would have mattered, and for whom. Drawing on philosophical work on contrastive and causal explanation, we diagnose three ways in which ML counterfactuals become ill-posed: they omit the relevant space of alternatives, hold causally dependent features fixed, and detach recourse from the agent and institution for whom it is supposed to be useful. We then give a target specification for trustworthy counterfactual explanation: a counterfactual contract that states the contrast, modal space, causal commitments, admissible actions, decision-rule stability assumptions, and intended answerability relation. Unlike prior critiques of actionability and causal recourse, we show that these are separable fields of well-posedness.