Symmetry-Constrained Gaussian Processes for Sample-Efficient Molecular Property Prediction
Kaustubh Bukkapatnam ⋅ Siddharth Karuturi ⋅ Laksh Patel
Abstract
Molecular property prediction from limited labeled data is a central bottleneck in computational chemistry and materials discovery. We introduce SYMGP, a Gaussian process framework whose kernel provably enforces the physical symmetries of molecular property functions: permutation invariance over atom indices, and invariance under the Euclidean group $E(3)$ of rotations, reflections, and translations. We show that symmetry-averaging reduces the effective dimension of the reproducing kernel Hilbert space, yielding tighter information-gain bounds and sublinear regret guarantees for the resulting active learning acquisition strategy. Concretely, we prove that the maximum information gain $\gamma_T$ for SYMGP with a squared-exponential base kernel scales as $\mathcal{O}((\log T)^{d_{\text{eff}}+1})$ where $d_{\text{eff}} < d_{\text{ambient}}$, giving a provably faster convergence rate than a symmetry-unaware baseline. Experiments on QM9 and FreeSolv demonstrate that SYMGP matches the accuracy of fully supervised deep models using up to $5\times$ fewer labeled examples, while producing well-calibrated predictive uncertainties that are required for closed-loop autonomous discovery pipelines.
Video
Chat is not available.
Successful Page Load