Dual Optimal Transport for Multi-Concept Composition: Structural Alignment and Texture Injection in Diffusion Models
Abstract
Lay Summary
People often want image generation models to place several specific subjects, such as a particular pet, toy, or accessory, into one coherent scene. This is difficult because the model must keep each subject recognizable, follow the text prompt, and avoid mixing the identities of different subjects. Existing methods often improve one of these goals at the cost of another, causing issues such as distorted layouts, lost details, or attributes leaking from one object to another. We propose OTComp, a training-free method for composing multiple reference concepts in text-to-image generation. Instead of retraining the model, OTComp guides the generation process in two steps: it first aligns the coarse structure of each concept with the intended scene layout, and then adds fine visual details such as textures and identity-specific features. This separation helps the model preserve both the overall scene and the details of each reference concept. Our work can support creative design, visual prototyping, and research on controllable image generation, while also highlighting the need for responsible use when generating images from personal or copyrighted references.