Neutral Reward Filtering for Fair Offline-to-Online Diffusion Alignment
Abstract
Preference optimization (PO) for text-to-image diffusion models can be viewed as an offline-to-online decision-making loop: a reward model learned from offline human preference data guides online adaptation of a generative policy. We show that CLIP-based rewards used in this loop encode perceived demographics in their visual features, so DPO/GRPO can amplify reward bias and narrow who appears in generated images. We propose Neutral Reward Filtering (NRF), a reward-side intervention that estimates demographic directions in the reward feature space with a one-time linear probe and removes them by orthogonal projection before rewards are computed. NRF is plug-and-play with both offline pairwise and online group-based PO, and requires no demographic labels during PO optimization beyond the audit stage. Experiments on SDXL and SD1.5 show that NRF restores most of the race/gender entropy lost during alignment while preserving cross-reward quality and prompt fidelity.