Tight Margin-Based Generalization Bounds for Voting Classifiers over Finite Hypothesis Sets
Abstract
We prove the first margin-based generalization bound for voting classifiers, that is asymptotically tight in the tradeoff between the size of the hypothesis set, the margin, the fraction of training points with the given margin, the number of training samples and the failure probability.
Lay Summary
If we assemble a group of machine learning models into a council and let them vote on a prediction, the council can perform significantly better than any individual member. The question is: how well can we expect the council to perform? In this work, we give mathematical guarantees for such a council vote. Based on the quality of the individual members and the data they are trained on, we can determine how well the final decision will perform on new, unseen data. Our result helps us understand how combining models improves performance both in practice and theory. Moreover, our guarantee is essentially the best possible of its kind.