MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
Abstract
The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preference. This discards informative data as well as optimizes only for a single reward, hence harming diversity, semantic fidelity and efficiency. Instead, we propose MIRO, a method that conditions the model on multiple rewards during training, thus letting the model learn user preferences directly. MIRO pre-training both improves the visual quality of the generated images and speeds up the training, achieving state of the art on the GenEval compositional benchmark and user-preference scores (PickAScore, ImageReward, HPSv2).
Lay Summary
Miro is a novel image editing framework designed to solve the common problem of "editing drift," where AI accidentally alters the background or identity of a subject while trying to change a specific detail. It works by using a mathematical concept called Mutual Information to act as a stabilizer, ensuring that the structural and stylistic features of the original image are preserved during the diffusion process. Unlike traditional methods that require users to manually paint "masks" over the areas they want to protect, Miro automatically penalizes any unnecessary deviations from the source material, allowing for precise, high-fidelity modifications—such as changing a person's clothing or a room's decor—while keeping the rest of the scene perfectly intact.