MICE-Bench: A Challenging and Comprehensive Benchmark for Multi-Reference Image Creation and Editing
Abstract
The paradigm of visual generation is rapidly shifting from single-image conditioning toward multi-image conditioning, making the ability to synthesize and edit images based on multiple visual references a critical capability. Despite this trend, existing benchmarks remain largely limited to single-reference scenarios or narrowly defined tasks, leaving model behavior under complex multi-concept composition insufficiently explored. To bridge this gap, we introduce MICE-Bench, a comprehensive benchmark for Multi-reference Image Creation and Editing. The benchmark is designed around three core principles: 1) heterogeneous concept composition across seven visual dimensions; 2) varying levels of constraint density, ranging from dual-concept to seven-concept configurations; 3) concept-centric data construction and benchmark evaluation, enabling fine-grained analysis of interactions among multiple concepts. MICE-Bench consists of 3,119 high-quality test cases within a unified concept space. Using an 8-dimensional evaluation metric, we systematically evaluate 13 state-of-the-art models. Our results show that although closed-source models maintain a clear performance advantage, all models experience notable degradation in concept consistency and physical realism as concept complexity increases.This indicates that current models rely on superficial composition rather than genuine multi-concept synthesis, highlighting substantial room for future improvement.
Lay Summary
AI models can now generate and edit images using multiple visual references simultaneously, such as combining a specific person, clothing item, and artistic style into one picture. However, there is no standardized method to evaluate how well models handle these complex combinations. We introduce MICE-Bench, a comprehensive framework featuring over 3,000 test cases. It challenges models to integrate up to seven distinct visual concepts—including identities, objects, and textures—and evaluates them on visual quality, instruction adherence, and physical realism. Testing 13 leading models reveals that while commercial systems outperform open-source alternatives, all suffer significant performance degradation as visual complexity increases. Models frequently lose details or violate physical laws, indicating they still struggle with genuine multi-concept synthesis. MICE-Bench exposes these current limitations and provides a clear roadmap for future improvements.