Robust Human-AI Complementarity under Uncertainty
Abstract
Machine learning models are often intended to augment rather than replace human decision-makers, by providing information that is complementary to human judgement. Yet, in practice, human decision makers routinely fail to realize such complementary gains, even when models provide useful signal. In this work, we study how asymmetric information about the quality of information available to a human decision maker vs. an AI impacts the ability of a decision maker to extract complementary value from AI predictions. We show that a key factor is the error correlation structure between human and AI predictions. In particular, when the AI's prediction errors are \textit{negatively correlated} with those of the human, the decision-maker can construct robust strategies which guarantee improvements in expected utility. We empirically investigate whether these conditions for complementarity arise in practice, using real-world forecasting benchmarks.
Lay Summary
AI is often meant to help people make better decisions, but in practice, people often fail to benefit from AI predictions even when those predictions contain useful information. This paper studies why that can happen when people are uncertain about how reliable an AI system is for a given task. We show that the key issue is not just whether the AI is accurate, but whether it makes different mistakes from the human. If AI errors tend to offset human errors, then a decision-maker can combine human and AI predictions in a way that reliably improves performance. But if the AI and human tend to make similar mistakes, then using the AI may offer little robust benefit, even when the AI contains useful information. Testing this theory on real-world forecasting tasks, we find that current language models often make errors that are positively correlated with human errors. Prompting strategies can reduce this overlap in some cases, but do not reliably eliminate it. These results suggest that AI systems meant to support human judgment should be trained and evaluated not only for accuracy, but for complementarity: providing information that humans are likely to miss.