SCOPE: Symbolic Constraint-Guided Active Perception
Abstract
Current multimodal reasoning systems often allocate extra inference-time computation to sampling additional textual rationales or tool-use trajectories. In visually ambiguous settings, however, the more consequential limitation is often the lack of an explicit intermediate state that records what the system currently believes about the scene, which alternatives remain plausible, which observations support those alternatives, and which new observations would be most informative for separating them. We introduce SCOPE, a symbolic activeperception framework that performs posterior sharpening over uncertainty-aware world hypotheses built from typed predicates, provenance, and explicit contradiction structure. SCOPE alternates between symbolic state induction, consistencybased world pruning, and conflict-driven evidence acquisition, so that additional computation is spent on refining explicit scene hypotheses rather than merely enlarging an unstructured pool of candidate outputs. This shift turns multimodal inference into a process of ambiguity management: the model preserves multiple plausible scene explanations early, prunes inconsistent ones as evidence accumulates, and focuses later perception on the few local conflicts that still influence the answer posterior. Across high-resolution and general multimodal reasoning benchmarks, SCOPE yields strong improvements over competitive baselines while also providing a more interpretable account of how extra computation changes the model’s internal decision state.