The Geometry of Reasoning Failure: Predicting Agent Errors from Semantic Scale Trajectories
Nikita Kazeev ⋅ Andrey Ustyuzhanin
Abstract
Selective prediction over agentic LLM outputs needs uncertainty signals. Direct examination of Chain-of-thought is a natural approach, but the reasoning text often becomes misleading and unreliable on the difficult task, where it's needed most. We introduce Semantic Level of Detail (SLoD), a continuous macro-to-micro abstraction axis over reasoning chunks, recoverable through three independent operationalizations (a SciBERT linear probe, a probe-free embedding-axis projection, and pairwise LLM judging), and use the shape of an agent's SLoD trajectory as a black-box selective-prediction signal. A LightGBM classifier over a stack of interpretable trajectory-shape detectors predicts SWE-agent failures at ROC-AUC $0.89\pm0.02$ and, used as an out-of-fold scorer to abstain from low-quality attempts, lifts Pass@1 on the FrontierScience Olympiad subset from $58.2\%$ to $68.0\%$ over five DeepSeek-V3.2 attempts (recovering $39.5\%$ of the Pass@5 oracle gap); the resulting coverage--accuracy curve dominates a length-only baseline at every operating point. The same trajectory features classify hallucination categories on the OpenManus subset of AgentHallu at macro-AUC $0.71$. We also report negative results on GPQA-Diamond's multiple-choice surface and on Qwen3-30B traces, showing that the signal is conditional on the generator and the problem.
Chat is not available.
Successful Page Load