Evaluator Failure Modes in Agentic Uncertainty Quantification
Suresh Raghu ⋅ Satwik Pandey ⋅ Shashwat Pandey
Abstract
Standard agentic UQ evaluations can hide trace-level failure modes. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized Trajectory Brier evaluate rankings, binwise calibration, or collapsed trajectory summaries, but they do not strictly elicit the prefix-conditioned success-probability process $q_t = \mathbb{P}^{\pi}(Y=1 \mid \mathcal{H}_t)$. This creates a practical diagnostic failure: an agent's confidence trace can appear acceptable under standard metrics while being badly mis-scaled as a probability stream for deferral, reflection, human handoff, or cost-weighted decisions. We characterize this failure mode theoretically and empirically. Theoretically, we show that Trajectory ECE is resolution-blind and that scalarized Trajectory Brier under common aggregators is not strictly proper for the trace. Empirically, we give two reproducible triggers. On Tau2-Bench, Platt recalibration of verbal confidence changes AUROC by only $\Delta/\mathrm{SE} \approx 0.3$, while changing a strictly proper trajectory score by $\Delta/\mathrm{SE} \approx 43$. On WebShop, complete-only evaluation drops $47.08\%$ of the assumption-valid working sample; the dropped max-step trajectories are roughly $3\times$ longer than completed ones, and censored-aware scoring changes the reported score. As a verified fix, we introduce the Trajectory Proper Score (TPS), a strictly proper trajectory-level evaluator built from any strictly proper binary score and positive trajectory weights, with a conditional-projection extension for administratively censored prefixes. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that evaluator choice can change empirical conclusions by margins far larger than bootstrap uncertainty.
Chat is not available.
Successful Page Load