Harmful Overfitting in Sobolev Spaces
Kedar Karhadkar ⋅ Alexander Sietsema ⋅ Deanna Needell ⋅ Guido Montufar
Abstract
Motivated by recent work on benign overfitting in overparameterized machine learning, we study the generalization behavior of functions in Sobolev spaces $W^{k, p}(\mathbb{R}^d)$ that perfectly fit a noisy training data set. Under assumptions of label noise and sufficient regularity in the data distribution, we show that approximately norm-minimizing interpolators, which are canonical solutions selected by smoothness bias, exhibit harmful overfitting: even as the training sample size $n \to \infty$, the generalization error remains bounded below by a positive constant with high probability. Our results hold for arbitrary values of $p \in [1, \infty)$, in contrast to prior results studying the Hilbert space case ($p = 2$) using kernel methods. Our proof uses a geometric argument which identifies harmful neighborhoods of the training data using Sobolev inequalities.
Lay Summary
Recent work has shown that machine learning models can perform surprisingly well at fitting data that they have not seen before. However, much of this work is specific to a high-dimensional setting where there is enough "room" to get things wrong without being penalized too much. We study models in a low-dimensional setting, where we prove that this over-performance does not occur. The reason for this difference is in the interaction between noise and smoothness. If a data point is mislabeled, then the model will do poorly not just on that data point, but also on a region around it. The level of harm from this region of misclassification is much worse in low dimensions than in high dimensions.
Successful Page Load