Deference-Aware Evaluation for Human-in-the-Loop AI Systems: A Unified Quality Signal for AI Operating Under Human Oversight
Abstract
Deference-Aware (DA) Evaluation is a framework for evaluating AI systems operating under human oversight. In human-in-the-loop (HITL) systems, standard evaluation metrics treat confident wrong answers and uncertain ones identically, creating systematic pressure toward overconfident AI, precisely the failure mode most dangerous when consequences are severe. DA evaluation credits appropriate deferral as correct within a single unified quality signal, with confident wrong answers as the only true failures. Because DA metrics are direct extensions of standard metrics on the same full-population denominator, the gap between them reads as a single-number diagnostic that distinguishes hidden calibration, genuine confident errors, and well-calibrated models. We instantiate the framework across five systematic-review datasets (2,729 studies) and six frontier models in evidence screening. The model that ranks first under standard recall-first evaluation is among the weakest once confident errors are counted: it reaches high recall through confident over-inclusion, with the highest confident error rate and almost no deferral. Under DA evaluation the best-calibrated model instead ranks first on every dimension (highest DA-Recall and DA-Specificity, lowest confident error and confident false-negative rates) and this ordering is stable across the full confidence-threshold range. We propose a minimum reporting standard for HITL AI with direct implications for procurement policy, regulatory frameworks, and DA-aware training. The framework extends naturally to agentic and multi-agent settings where confidence-triggered fallback is the emerging deployment pattern.