Learning $U$-Statistics with Active Inference
Abstract
Lay Summary
Many statistical studies rely on U-statistics, a broad class of tools used to measure quantities such as income inequality, ranking performance, or dependence between variables. For example, the Gini index is a U-statistic commonly used to measure income inequality. In many modern applications, however, the labels or outcomes needed to compute these quantities are expensive or time-consuming to obtain. This paper studies how to estimate U-statistics accurately under a limited labeling budget. Our method uses machine learning predictions in two complementary ways: it first builds a prediction-based estimate from all available data, and then uses the predictions to select the most informative data points for labeling. By combining predictions with actively collected labels, our method obtains more reliable estimates using fewer labels and produces uncertainty ranges for the final estimate.