NAVIGATE: Evaluating Visual-Guided Search Decision-Making on the Open Web
Abstract
Vision–Language Models (VLMs) are increasingly deployed with web search tools, yet we still lack benchmarks that isolate a critical capability for real-world use: deciding when to search and how to steer search from ambiguous visual evidence, especially when multiple images provide overlapping or conflicting cues. We introduce NAVIGATE, a novel benchmark centered on images as primary evidence for open-web search planning and multi-step reasoning. It contains 500 questions across 20 domains and spans three difficulty tiers, from single-image, self-contained problems to multi-image joint search and multi-domain composition. Unlike prior benchmarks that specify explicit search targets, NAVIGATE evaluates search decision-making: models must infer whether external search is necessary and iteratively refine search directions based on holistic reasoning over visual cues. Across a broad set of VLMs and search-enabled systems, performance remains low, Gemini-3-Pro-Preview-Search reaches only 36.4% accuracy, highlighting persistent failures in cross-image grounding, search triggering, and search strategy coordination. We will release NAVIGATE publicly.
Lay Summary
Imagine showing an AI a photo of a rare vintage car and a mysterious landmark in the background, then asking where it was taken. While humans would naturally know to search for the landmark to identify the location, current AI often gets "confused" or tries to guess without enough evidence. To bridge this gap, we created NAVIGATE, a benchmark that tests if AI can think like a detective. Instead of just "identifying" objects, AI must now decide when to search the open web and how to navigate through complex visual clues across 20 different fields, from art to technology. Our findings show a major gap in current technology: even top-tier models fail over 60% of these challenges. NAVIGATE sets a new standard for the next generation of AI, moving us closer to digital assistants that don't just "see," but actively and intelligently "research" to solve the complex problems we face in the real world.