Reward-Aligning Few-Step Flow Models with Integrated Regularizers
Abstract
Flow-based generative models are increasingly adapted to downstream rewards, preferences, and design objectives. Most existing approaches fine-tune the time-dependent drift of the underlying continuous flow. This is costly and poorly matched to the growing use of accelerated one-step and few-step samplers. In this work, we instead align the accelerated sampler directly by fine-tuning a flow map. Its endpoint outputs define an adapted terminal distribution on which rewards can be evaluated, providing a strong training signal to the fast generator. To regularize this sampler, we use the same terminal samples to construct an interpolant-based family of intermediate marginals, and penalize their deviation from the base model marginals across time using integrated divergences. By training a single flow map, we avoid requiring separate generator and score networks, and remove the need for expensive rollouts during adaptation. Experiments on ImageNet show that our method effectively improves reward-aligned sample quality while retaining the efficiency of few-step generation.