Reflection Anchors for Interpretable Compositional Visual Reasoning in Multimodal Reinforcement Learning
Abstract
Long-chain reasoning in vision--language models can be viewed as a composition of intermediate operations: inspecting visual evidence, relating entities, revising hypotheses, and committing to an answer. A common failure mode is that this composition becomes increasingly textual, with later operations losing dependence on the image. We study this as a problem of compositional visual retention in multimodal reinforcement learning. We propose RAPO, a GRPO-based method that updates sparse reflection anchors: high-entropy positions where the model is uncertain about the next reasoning operation. RAPO optimizes a chain-masked finite-window KL objective that encourages visual evidence to propagate through downstream reasoning, without architectural changes or inference-time visual re-injection. Experiments across reasoning-intensive and general-domain multimodal benchmarks show consistent gains over strong RL and visual-grounding baselines. Mechanism analyses further show that anchors align with interpretable operation-selecting tokens and that visual influence from these anchors persists into later reasoning. These results suggest that sparse, interpretable policy interventions can make long-chain multimodal reasoning more compositionally grounded.