Embodied Task Planning via Graph-Informed Action Generation with Large Language Models
Abstract
While Large Language Models (LLMs) have demonstrated strong zero-shot reasoning capabilities, their deployment as embodied agents still faces fundamental challenges in long-horizon planning. Unlike open-ended text generation, embodied agents must decompose high-level intents into actionable sub-goals while adhering to the constraints of a dynamic environment. Standard LLM planners frequently fail to maintain strategy coherence over extended horizons due to context window limitations or hallucinate state transitions that violate environment constraints. We propose GiG, a planning framework that structures embodied agents’ memory using a Graph-in-Graph architecture. Our approach employs a Graph Neural Network (GNN) to encode environmental states into embeddings, organizing these embeddings into action-connected execution trace graphs within an experience memory bank. GiG enables retrieval of structurally-similar priors, allowing agents to ground current decisions in relevant past structural patterns. Furthermore, we introduce a bounded lookahead module that leverages symbolic transition logic to enhance the agent’s planning capabilities through grounded action projections. We evaluate our framework on three embodied planning benchmarks—Robotouille Synchronous, Robotouille Asynchronous, and ALFWorld. Our method outperforms state-of-the-art baselines, achieving Pass@1 performance gains of up to 22% on Robotouille Synchronous, 37% on Asynchronous, and 15% on ALFWorld while maintaining comparable or lower computational cost.
Lay Summary
LLMs are becoming remarkably good at generating text, but they still struggle to compose coherent plans. In complex environments, these AI agents often get "lost," failing to remember their goals or becoming stuck in repetitive cycles because they cannot effectively track long-term tasks or environmental constraints. To address this, we introduced a new framework called GiG (Graph-in-Graph). Instead of relying on a simple, growing list of past actions, which often confuses AI models, our system organizes the agent’s environment and history into a structured, two-tier memory graph. This memory structure acts like a "roadmap," providing the agent with an organized understanding of its surroundings and allowing it to retrieve successful strategies from past experiences. Our method significantly improves planning performance on complex tasks while remaining computationally efficient.