IGG: A Benchmark for Interactive GUI Grounding under Visibility Constraints
Abstract
GUI grounding benchmarks evaluate whether a VLM agent can localize a target element from a single static screenshot, typically assuming that the target is already visible. Real-world interfaces, however, often involve visibility constraints such as off-screen targets, hover-based descriptions, occlusions, and delayed activation, so a correct click first requires an interaction that reveals the target. In these settings, grounding requires not only visual matching but also interaction to recover target visibility. We introduce Interactive GUI Grounding (IGG), a benchmark for grounding under limited observability. In IGG, the target is not directly localizable or actionable from the initial screenshot, and agents must expose the target before localization. We define a minimal action space for visibility recovery and a three-level taxonomy of GUI constraints spanning single-state, multi-state, and temporal and advanced settings, with seven sub-types, enabling systematic evaluation of GUI agents under diverse visibility constraints.