LOGIV: Logic-Graph Inference with VAL-Verification for Long-Horizon Robotic Manipulation
Abstract
Long-term robotic operation has long been plagued by temporal failures during execution, as static task instructions (i.e., language conditions) fail to provide dynamic guidance for complex stages. This results in stage confusion and goal deviation in multi-stage scenarios, despite the strong short-term reaction capabilities of the underlying actuators. In this paper, we demonstrate that achieving robust long-term autonomy does not require retraining the actuators or complex hierarchical architectures, but rather temporal reparameterization of the task-instruction interface. We propose a training-free dual-system reasoning framework, LOGIV (LOgic-Graph Inference with VAL-verification), which decomposes global instructions into dynamic, multi-stage natural language priors. To ensure logical consistency, we introduce a graph-based self-correction mechanism that utilizes formal verification and standardized repair operators to autonomously correct plans generated by large language models. Experimental results on datasets such as DROID, AgiBot, EgoDex, and RoboTwin 2.0 show that our framework significantly improves task success rates without modifying the underlying model weights, with more pronounced effects on more complex tasks. Due to its architecture-agnostic and lightweight nature, LOGIV offers an efficient and plug-and-play solution for bridging the long-term temporal gap in existing world action models.