Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation
Abstract
Lay Summary
Although many text-to-image systems can generate high-quality images, they still struggle with complex prompts involving multiple attributes, such as object counts, colors, or spatial arrangements. Even when a system detects visual discrepancies in a suboptimal image, it often fails to provide actionable instructions for correction. This paper investigates the gap between visual error detection and subsequent rectification. We introduce R3-Bench, a benchmark comprising 670 expertly curated instances that evaluates the capability of models to verify image-text alignment, explain discrepancies, and propose revision strategies. Furthermore, we propose R3-Refiner, a training methodology designed to translate visual reasoning capabilities into explicit and actionable rectification instructions. Extensive experiments demonstrate that our approach significantly enhances both error diagnosis and image correction while generalizing across various image generation frameworks. These advancements enhance the reliability of text-to-image tools for diverse applications, though deployment must be accompanied by robust safeguards against misleading synthetic content.