GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance
Abstract
Guidance methods, such as classifier-free guidance (CFG) and auto-guidance (AG), have advanced noise-to-data generation in diffusion models. Recently, bridge models have introduced a data-to-data generative process that can exploit an instructive clean prior. In this work, inspired by previous methods creating quality difference between denoising results as guidance, we propose a training-free bridge guidance method, termed Prior Guidance (PG). Specifically, we introduce a weak prior, which is unseen during bridge pre-training, hindering prior exploitation and thereby degrading denoising result. Then, we contrast it with the seen prior to highlight and enhance prior exploitation via a scaling factor. Moreover, we analyze the underlying mechanism of prior exploitation in the bridge process and design frequency-modulated prior guidance (FMPG), which tailors the guidance scale to low- and high-frequency bands coherent with bridge generative dynamics. To address prior exploitation in image in-painting, we develop a cascaded framework, CFG-FMPG, which first generates a noisy hidden representation via CFG and then exploits it as a generative prior with FMPG, fulfilling their complementary strengths without compromising inference efficiency. Experiments demonstrate that our PG methods consistently improve pre-trained bridge models across diverse image translation tasks.
Lay Summary
Guidance methods have enhanced condition alignment or score accuracy of diffusion generation with Classifier-Free guidance (CFG) or Auto-Guidance (AG), respectively, through contrasting two denoising terms with quality difference. Inspired by their mechanisms, we design a training-free guidance method for bridge models, termed Prior Guidance (PG). Bridge models develop a data-to-data generative process that can directly exploit a clean and informative prior in sampling, overcoming the restriction of noisy prior in diffusion models. Hence, we further amplify this advantage with guidance methods in inference. Specifically, we manually construct a low-quality prior representation which is unseen during bridge pre-training, hindering prior exploitation and degrading bridge denoising quality. By contrasting this low-quality term and the denoising result from seen prior, we emphasize prior exploitation and improve bridge generation results. Furthermore, we investigate the prior exploitation mechanism in bridge generative process, designing frequency-modulated prior guidance (FMPG) and a cascaded guidance framework, CFG-FMPG. These innovations further improve the efficiency of PG and combine its advantage with CFG without compromising sampling efficiency. Extensive experiments verify that PG methods enhance pretrained bridge models, DDBM and DBIM, across diverse image translation tasks.