Mechanistic Interpretability of Adversarial Suffixes Reveals Non-Robust Shortcuts of Safety Monitors
Abstract
We apply mechanistic interpretability to adversarial robustness, using causal interventions to reverse-engineer how Greedy Coordinate Gradient (GCG) attacks fool safety guardrails. Activation and attribution patching across four BERT-family classifiers reveals a consistent two-stage circuit: adversarial signal forms as a payload in early-layer MLP activations at suffix positions, then routes to non-suffix positions via attention keys at suffix positions. While prior interpretability work has centered on attention hijacking to study the success of adversarial suffix attacks, we find that GCG converges on a sparse, interpretable set of MLP neurons across independently optimized attacks. We use these neurons' interpretations to construct human readable adversarial examples that exploit non-robust features that transfer to NemoGuard safety monitors an order of magnitude larger.