TRACE: Why LLM Agents Fail in Multi-Step and Multi-Turn Environments?
Aadit Roy Chowdhury ⋅ Emre Can Acikgoz ⋅ Cheng Qian ⋅ Dilek Hakkani-Tür ⋅ Gokhan Tur
Abstract
LLM-based agents have made remarkable progress in performing complex tasks across multi-turn interactions and in environments that require multi-step reasoning. Although recent advances have enabled more dynamic interactions with users and environments, these models still fail in ways that are not well defined and understood. In this paper, we discuss what agent abilities are lacking when LLMs fail in multi-turn settings. We perform a comprehensive analysis of around 10,000 trajectories spanning four popular agentic benchmarks: $\tau^2$-bench, BFCL v4, ALFWorld, and ACEBench and seven LLMs covering four closed-source models: GPT-5.4 Mini, DeepSeek-R1, Kimi-K2.5, and Llama-3.3-70B-Instruct-Turbo and three open-source models at different scales: Qwen3 (1.7B, 4B, 8B). We introduce TRACE, a capability-based taxonomy that organizes failures into three categories: A) action–plan failures, B) monitoring failures, and C) outcome evaluation failures. Each failure type is mapped to specific agent skills, such as self-awareness, self-correction, proactivity, state-tracking, and policy adherence. We show that certain failure patterns recur across benchmarks and model families and can be attributed to poor agent capabilities. Our analysis introduces a principled taxonomy of the failure patterns and highlights the need for training methods that provide fine-grained, step-level feedback to LLM agents to improve these capabilities in multi-turn and multi-step reasoning environments.
Chat is not available.
Successful Page Load