Compositional Investigation: Why Reasoning Enables Tool-Using Agents to Fix What They Diagnose
Dhatri C ⋅ Tadisetty S Yashwanth
Abstract
We introduce rca llms, a closed-loop bench- mark in which an agent must diagnose and patch microservice bugs by composing four sub-skills (evidence gathering, cross-service correlation, code localization, and code editing), with cor- rectness verified by a deterministic reproducer rather than an LLM judge. Across six frontier models and five multi-hop incidents, all three rea- soning variants achieve 5/5 pass rates while all three non-reasoning variants score 1/5, with most failures being non-terminations. The same base model (grok-4-1-fast) scores 5/5 with rea- soning and 1/5 without, isolating chain-of-thought as a compositional controller that commits to sub- conclusions and advances the pipeline, rather than a simple accuracy booster.
Chat is not available.
Successful Page Load