Test-Time Guidance for Flow-Based Generative Models via Parallel Tempering on Source Distributions
Abstract
Generative models that transport a simple source distribution to a complex data distribution—such as diffusion and flow-based models—are central to high‑fidelity data generation. Test-time guidance can further steer pretrained models toward user-specified high-reward regions without costly retraining. However, existing guidance methods face critical limitations: they struggle with non-differentiable rewards, fail to navigate complex landscapes, and often lack theoretical guarantees on generation performance. We propose {\it Source Parallel Tempering (SPT)}, a gradient‑free test‑time guidance framework that operates entirely in source space, leveraging its simpler geometry to avoid the complexities of the data manifold. SPT couples a local exploration kernel with parallel tempering, enabling efficient barrier crossing and robust discovery of high‑reward modes. Theoretically, we provide a new error bound linking training-time approximation error to test-time guidance performance.Empirically, SPT significantly improves over state-of-the-art methods on benchmark tasks in conditional image synthesis and dynamical system trajectory sampling.
Lay Summary
Generative models are increasingly used in scientific and creative applications where outputs must satisfy specific objectives. For example, a user may want to generate images with a desired aesthetic or trajectories satisfying physical constraints. Guiding pretrained models toward such objectives without retraining remains challenging, especially when the reward function is black-box or non-differentiable. We introduce Source Parallel Tempering (SPT), a gradient-free test-time guidance method that operates in the model’s latent source space rather than the complex data space. This simpler geometry makes exploration more efficient and robust. SPT combines local random exploration with multiple parallel search processes at different “temperatures,” allowing it to escape poor local optima and discover high-reward solutions. SPT consistently outperforms existing guidance methods on image generation and dynamical system trajectory tasks, particularly in settings with difficult black-box rewards. We also provide theoretical guarantees for SPT that link generative model approximation quality to guided generation performance, giving practitioners a principled understanding of when and why this test-time guidance method succeeds.