Coarse-Grained Boltzmann Generators
Abstract
Sampling equilibrium molecular configurations from the Boltzmann distribution is a longstanding challenge. Boltzmann Generators (BGs) address this by combining exact-likelihood generative models with importance sampling, but practical scalability is limited. Meanwhile, coarse-grained surrogates enable the modeling of larger systems by reducing effective dimensionality, yet often lack a reweighting procedure required to ensure asymptotically correct statistics. In this work, we propose Coarse-Grained Boltzmann Generators (CG-BGs), a framework for reduced-order generative modeling with importance sampling in coarse-grained coordinate space. CG-BGs generate samples using a flow-based model and reweight them using a learned potential of mean force (PMF). We show that the PMF can be learned from rapidly converged trajectories via enhanced sampling force matching. Experiments demonstrate that CG-BGs capture solvent-mediated interactions in highly reduced representations while substantially reducing computational cost relative to atomistic BGs, providing a practical route toward equilibrium sampling of larger molecular systems.
Lay Summary
Molecular simulations are widely used to study how molecules move and interact, but they become very slow for large or complex systems. This is because accurate simulations must follow the motion of every atom over time, which is computationally expensive, and they can struggle to capture rare but important molecular states. We propose a machine learning approach that reduces this cost by representing molecules in a simplified form instead of tracking all atoms directly. A generative model produces new molecular configurations in this reduced space, and a physics-based correction step adjusts these samples so that they obey the correct statistical laws. This correction is based on a learned energy function that can be trained efficiently from shorter, partially guided simulation data. Our results show that this approach can reproduce accurate molecular statistics while requiring significantly less computation than standard simulation methods. It also provides a viable way to scale machine learning-based molecular sampling to larger systems.