Efficient Training of Boltzmann Generators Using Off-Policy Log-Dispersion Regularization
Abstract
Sampling from unnormalized probability densities is a central challenge in computational science. Boltzmann generators are generative models that enable independent sampling from the Boltzmann distribution of physical systems at a given temperature. However, their practical success depends on data-efficient training, as both simulation data and target energy evaluations are costly. To this end, we propose off-policy log-dispersion regularization (LDR), a novel regularization framework that builds on a generalization of the log-variance objective. We apply LDR in the off-policy setting in combination with standard data-based training objectives, without requiring additional on-policy samples. LDR acts as a shape regularizer of the energy landscape by leveraging additional information in the form of target energy labels. The proposed regularization framework is broadly applicable, supporting unbiased or biased simulation datasets as well as purely variational training without access to target samples. Across all benchmarks, LDR improves both final performance and data efficiency, with sample efficiency gains of up to one order of magnitude.
Lay Summary
Physical simulations often require long, costly trajectories to generate useful molecular samples. Boltzmann generators aim to replace this with direct, one-shot sample generation leveraging generative models, but training them can itself require large amounts of expensive simulation data. We introduce a way to use energy information that is often already available but typically not used for training, as an extra training signal. This makes Boltzmann generators substantially more data-efficient, reducing the cost of learning accurate molecular samplers.