Prediction-Powered Inference Across Many Tasks for AI Evaluation
Abstract
Many applications require statistically valid inference across many related tasks, while providing only a handful of high-quality labels per task. In AI evaluation and testing, these tasks may correspond to model behaviors across prompts, subgroups, or hypotheses; in social science surveys, they may correspond to related questions, populations, or measurement conditions. Prediction-powered inference uses inexpensive proxy measurements to improve inference from limited labels, but standard methods operate task-by-task and therefore struggle in the small-label regime. We introduce a multi-task prediction-powered inference framework that borrows strength across tasks without pooling away validity. Our methods recalibrate surrogate outcomes using labeled data from other tasks while retaining valid task-level rectification and empirically tighter per-task confidence intervals. We prove that power gains beyond power-tuned PPI require nonlinear structure in the proxy–ground-truth relationship and that affine cross-task recalibrations are oracle-equivalent to using the original proxy. We complement our theoretical findings with experiments on semi-synthetic datasets and a case study auditing language models on election-related information during the 2024 U.S. presidential election. Using a large human-annotation study, we show that cross-task surrogate learning can substantially reduce confidence interval widths when labels are scarce, and increase the power of corresponding hypothesis tests.