Diagnosing Failure Modes in MCP Tool-Using Agents: A Stage-Aware Taxonomy and Evidence-Gap Interventions
Abstract
Tool-using agents are often assumed to fail because they select the wrong tools. We show that this explanation is insufficient in realistic MCP environments. We introduce a stage-aware failure taxonomy that assigns failed trajectories to No-tool, Routing, Workflow, Request/Parameter, Evidence Use, or Runtime failures, separating tool-selection errors from downstream execution failures. Across five agents on 427 MCP-Atlas tasks, restricting agents to gold tools yields only modest gains and leaves many failures in workflow continuation, request grounding, evidence use, and runtime stability. We then use evidence-gap and tool-path guidance as diagnostic probes. Proactive guidance is more effective than reactive post-final feedback, but many failures remain or shift to downstream stages. A trace-level case study further shows that agents can use the same relevant tool families yet diverge depending on query precision, evidence linking, and workflow persistence. Our results indicate that gold-tool access is necessary but not sufficient: reliable MCP agents require evidence-state tracking, query repair, and workflow-level verification beyond better tool selection.