Symmetries in PAC-Bayesian Learning
Abstract
Symmetries are known to improve the empirical performance of machine learning models, yet theoretical guarantees explaining these gains remain limited. Prior work has focused mainly on compact group symmetries and often assumes that the data distribution itself is invariant, an assumption rarely satisfied in real-world applications. In this work, we extend generalization guarantees to the broader setting of non-compact symmetries, such as translations and to non-invariant data distributions. Building on the PAC-Bayes framework, we adapt and tighten existing bounds, demonstrating the approach on McAllester's PAC-Bayes bound while showing that it applies to a wide range of PAC-Bayes bounds. We validate our theory with experiments on several datasets with non-uniform and non-compact transformations, where the derived guarantees not only hold but also improve upon prior results. These findings provide theoretical evidence that, for symmetric data, symmetric models are preferable beyond the narrow setting of compact groups and invariant distributions, opening the way to a more general understanding of symmetries in machine learning.
Lay Summary
Machine learning models often perform better when they account for symmetries in data for example, recognizing that an object remains the same even if it is shifted or rotated. Although such symmetry-aware models are widely used in practice, existing mathematical guarantees apply only to highly idealized settings with special symmetries and data distributions. In this work, we develop new theoretical guarantees for a much broader and more realistic setting. Our analysis covers non-compact symmetries, such as arbitrary translations, and also applies when the data distribution itself is not symmetric. Using the PAC-Bayes framework, we derive tighter generalization guarantees. We validate the theory on several data sets, demonstrating that the guarantees hold in practice. Our findings provide theoretical support for designing symmetry-aware machine learning systems in real-world applications, including robotics, medical imaging, and scientific data analysis.