WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Abstract
Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and world knowledge integration in text-to-image generation. To address this challenge, we propose WISE, the first benchmark specifically designed for World Knowledge-Informed Semantic Evaluation. WISE moves beyond simple word-pixel mapping by challenging models with 1000 meticulously crafted prompts across 25 subdomains in cultural common sense, spatio-temporal reasoning, and natural science. To overcome the limitations of traditional CLIP metric, we introduce WiScore, a novel quantitative metric for assessing knowledge-image alignment. Through comprehensive testing of 20 models (10 dedicated T2I models and 10 unified multimodal models) using 1,000 structured prompts spanning 25 subdomains, our findings reveal significant limitations in their ability to effectively integrate and apply world knowledge during image generation, highlighting critical pathways for enhancing knowledge incorporation and application in next-generation T2I models. Code and data will be available.
Lay Summary
AI systems can now create impressive images from text, but they often fail when a prompt requires real-world knowledge rather than simply matching words to visible objects. For example, generating “Einstein’s favorite musical instrument” requires knowing that Einstein played the violin, not just drawing the words in the prompt. In this paper, we introduce WISE, a new test for evaluating whether image-generation models can use such world knowledge when creating images. WISE contains 1,000 carefully designed prompts covering cultural common sense, time and space reasoning, and natural science. We use WISE to evaluate 20 image-generation models, including both standard image generators and newer multimodal models that combine language and vision abilities. Our results show that current models still struggle to turn knowledge and reasoning into accurate images. This suggests that future image-generation systems need stronger connections between understanding the world and visually depicting it.