VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
Abstract
Vision-language models (VLMs) perform strongly on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, where a model must follow a selected path through successive local continuations. To isolate this ability, we design controlled tasks with nearby competitors while minimizing semantic and topological ambiguity such as crossings and overlaps. Across tasks, even state-of-the-art VLMs often lose the target path and switch to nearby alternatives, especially when distractors are locally similar. Behavioral and internal analyses show that these failures stem from local competition: nearby similar distractors pull the model away from the true continuation. Standard remedies offer limited relief, as scaling, reasoning, and explicit tracing instructions fail to restore stable path following. Finally, tests on tangled cables and metro maps show that the same path-switching failure persists beyond controlled settings.