Denser Input, Worse Agent: Compound Bottlenecks and a Capability Threshold in VLM Spatial Reasoning
Abstract
Vision–language models (VLMs) lag text-only LLMs on active spatial reasoning, even though the visual input carries strictly denser information than the matching textual scene description (Zhang et al., 2026a). We trace this persistent gap to two compounding bottlenecks: (B1) a VLM-specific perception bottleneck concentrated in object orientation, and (B2) a universal belief-instability bottleneck — both LLM and VLM drift at comparable rates across turns. We attempt to close this gap with a multi-agent pipeline of perception sub-agents and an external cognitive map. Across three frontier models, the pipeline reliably improves cogmap quality (the agent’s internal spatial belief of the environment), but these representational gains translate to evaluation accuracy (downstream task performance) only above a model-capability threshold — yielding pipeline-helps, pipeline-neutral, and pipeline-hurts outcomes. We trace this divergence to a single dimension: world coverage (the fraction of world objects placed in the cogmap, driven by the agent’s exploration policy).