FrontierSpatial: Holistic Evaluation of Spatial Reasoning in Multimodal Large Language Models
Abstract
Spatial reasoning is a core capability for intelligent systems, yet evaluation of multimodal large language models (MLLMs) remains fragmented across narrowly focused benchmarks targeting isolated skills. To address this limitation, we introduce FrontierSpatial, the first unified large-scale benchmark and evaluation harness for holistic spatial reasoning in multimodal models. FrontierSpatial consolidates and extends 23 datasets into 7,838 questions and 11,825 images, organized within a four-level hierarchy spanning 7 categories, 19 sub-categories, and 65 fine-grained tasks, including Physical, Geometric, Knowledge-Grounded, Relational/Causal, Temporal, Quantitative, and Perceptual Reasoning. Beyond the benchmark, we contribute: (1) an automated curation pipeline that distills over 60,000 raw instances into a balanced, high-quality benchmark; (2) an automated evaluation harness for reproducible assessment of 16 state-of-the-art MLLMs; and (3) an automated error-analysis framework that generates structured diagnostics across the task hierarchy. Experiments reveal large performance disparities across spatial reasoning categories and systematic weaknesses that remain hidden when benchmarks are studied independently. FrontierSpatial provides a rigorous and extensible foundation for holistic multimodal spatial reasoning research.