Before the Fall: Delta Minimal Failing Prefixes for Local Tool-Use Agent Failures
Abstract
Final success rates conflate qualitatively different failures in tool-use agents: a model may fail from the initial interface, enter a mid-trajectory cascade, behave unstably under replay, or make an irreversible state change. We introduce ∆-Minimal Failing Prefixes (∆-MFP), a counterfactual replay diagnostic that compares failure probability from a trajectory prefix to failure probability from the initial state, separating prefix-0 capability failures from nontrivial failure basins. We build a local, stateful tool-use suite with 120 calendar, refund/order, file/email, and inventory tasks, evaluate local Qwen2.5 and Llama-3.1 agents under four tool interfaces, and run three failure probes using local single-GPU inference: natural failures from calibrated cells, persistent fault injection as a positive-control verification of the replay engine, and soft fault injection that perturbs evidence, memory, or arguments without committing a state violation. On 25 natural failed traces, ∆-MFP identifies 13 nontrivial failure basins, 5 prefix-0 failures, and 7 unstable traces. On 50 retained soft-fault traces across six subtypes, the replay phase diagram finds 7 nontrivial ∆-MFPs, 23 unstable traces, and 20 prefix-0 failures, showing that replay-based localization should report uncertainty rather than force every failure into a root-cause prefix. We further probe lightweight repairs and a rule-based router, reporting them as low-N diagnostics rather than method rankings. Overall, ∆-MFP provides a reproducible failure-regime profile for local tool-use agents rather than a single aggregate success metric.