GenExam: A Multidisciplinary Text-to-Image Exam
Abstract
Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style benchmarks mainly focus on understanding and reasoning tasks, and current generation benchmarks emphasize the illustration of world knowledge and visual concepts, neglecting the evaluation of rigorous drawing exams. We introduce GenExam, the first benchmark for multidisciplinary text-to-image exams, featuring 1,000 samples across 10 subjects with exam-style prompts organized under a four-level taxonomy. Each problem is equipped with ground-truth images and fine-grained scoring points to enable a precise evaluation of semantic correctness and visual plausibility. Experiments on 17 text-to-image and unified models demonstrate the great challenge of GenExam and the huge gap where open-source models consistently lag behind the leading closed-source ones. By framing image generation as an exam, GenExam offers a rigorous assessment of models' ability to integrate understanding, reasoning, and generation, providing insights on the path to intelligent generative models. Our benchmark and evaluation code will be released.
Lay Summary
Exams are one of the clearest ways to measure human expertise because they require understanding, reasoning, and the ability to produce correct answers. However, most existing AI benchmarks for image generation focus on creating visually appealing pictures rather than solving rigorous exam-style drawing problems. As a result, it remains unclear whether modern generative models can truly combine reasoning with image creation. To address this gap, we introduce GenExam, the first benchmark designed to test text-to-image models with multidisciplinary exam questions. The benchmark contains 1,000 problems from 10 subjects, including tasks that require accurate diagrams, scientific illustrations, and structured visual reasoning. Each question includes a reference image and detailed scoring criteria so models can be evaluated not only on visual quality, but also on semantic correctness and factual accuracy. We evaluated 17 state-of-the-art image generation and multimodal models on GenExam. Our results show that these tasks are highly challenging, and that open-source models still lag significantly behind leading closed-source systems. By treating image generation as an exam, GenExam provides a more rigorous way to measure whether AI systems can integrate understanding, reasoning, and generation, helping researchers better track progress toward more capable and reliable generative models.