Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM Gameplay
Abstract
Lay Summary
AI systems that process both images and text can describe photos and answer visual questions, but they struggle with tasks that require understanding rules and reasoning about what they see. When these systems play board games like chess, they make mistakes, but until now it was unclear whether they fail because they cannot see the board correctly, cannot understand the rules, or cannot think ahead. We designed controlled tests to separate these abilities. Surprisingly, when these systems misread a game board, the errors are not random: neighboring pieces tend to be shifted together in the same direction, as if entire regions of the image slide as a block. We also found that even when the systems read the board perfectly, they still fail to apply game rules correctly, and that they can verify whether a completed move was legal much more easily than they can predict whether a future move is possible. Using more powerful versions of these systems or giving them step-by-step instructions does not fix these problems. These findings matter beyond games. Any application where AI needs to look at structured visual information and reason about it, from reading documents to guiding robots, may face the same blind spots.