Spiking the training data to adjust for test set contamination
Abstract
Statistical adjustment for test set contamination is underexplored. Our core proposal is to \textit{spike} the training data and intentionally insert test examples at known rates. The spiked examples can then be used to calibrate memorization detectors and enable principled statistical adjustment. To evaluate adjustment estimators, we first present a simulation framework based on the Hubble models. Hubble models come in minimal pairs, where the perturbed model was intentionally contaminated with several test sets, while the standard model was not, serving as the counterfactual and adjustment target. By sampling contaminated test sets, we can evaluate adjustment estimators that use information from a memorization predictor, a correctness predictor, or both. We establish basic statistical intuitions and show that estimators leveraging memorization and correctness information perform better across a wide range of operating settings over the naive estimator which makes no adjustment. We then instantiate and evaluate several memorization and correctness predictors in the simulation, and find that simple predictors such as applying Platt scaling to membership inference methods provide good signal for adjustment. Finally, we examine the practical considerations of spiking. Simple memorization predictors typically need no more than 10 examples for calibration and decently transfer from one dataset to another. Taken together, spiking is a promising solution to test set contamination.