The Probe Detected the Banner, Not the Attack: Trace Diagnostics for an Agentic-AI Evaluation Failure
Abstract
Hidden-state probes are increasingly used as pre-output monitors for multimodal computer-use agents under indirect prompt injection, but a high clean-vs-attack AUC may certify the wrong property. We study this metric-interpretation failure in a teacher-forced Qwen2.5-VL-7B / Mind2Web replay setting: a linear probe reaches AUC 0.998 on a visible-overlay IPI surface, a result that could be misread as evidence that the agent internally detects malicious instructions. We show that this interpretation is not licensed by the evaluation label, which only marks injection-surface presence. Two same-trace diagnostics expose the gap. First, a four-scalar metadata baseline saturates on text-side DOM/tool injections, revealing paired-construction leakage. Second, visually matched overlay controls show that a scrambled, non-semantic banner can match the malicious overlay’s clean-vs-attack AUC, while direct malicious-vs-control AUC fails to reliably separate the two. These controls do not prove content blindness, but they invalidate an unqualified semantic interpretation of the headline AUC. We package the diagnostics as an evaluation-level intervention: a claim-boundary checklist specifying what a high probing AUC can and cannot support. The result is a carefully scoped negative case study of an agentic-AI evaluation failure, not a deployed defense.