Trace-Level Failure Boundaries for Engineering Agents
Abstract
Verifier-augmented agents can look reliable in clean runs while failing at the boundary where tool outputs, artifact fields, and later checks diverge. We study this problem in SysML v2 link-budget generation and introduce a replayable trace-level protocol for engineering agents: deterministic labels for verifier invocation and result use, fixed gates over emitted artifacts, and tool-free repair tests on frozen failures. The simple faithful-tool calibration is deliberately treated as weak evidence because the exposed tool and the final verifier share the same physics. The boundary probes are the main object: they show failures from schema ambiguity, audit-surface mismatch, and stale artifact memory, and they test whether explicit verifier-visible contracts or repair signals move those failures. The strongest repair lesson is negative as well as positive: more verifier text is not automatically better, while structured quarantine of suspect memory is a useful but bounded intervention. The contribution is not a higher pass rate, but a reproducible way to localize when verifier use stops being faithful and which repair signals actually change the failed artifact.