Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models
Arth Ashish Singh
Abstract
Safety alignment in diffusion language models (dLLMs) relies heavily on a trajectory-level invariant: that committed tokens are permanent. We show that violating this invariant, by re-masking committed refusal tokens and injecting a short affirmative prefix, achieves 74–82% Attack Success Rate (ASR) on HarmBench across all three publicly available safety-tuned dLLMs, rising to 92–98% with a generic 8-token compliance prefix. We call this attack **TrajHijack**; to our knowledge, it is the first *gradient-free monotonicity-violating* trajectory-level attack on dLLMs (concurrent priming/anchoring work is trajectory-level but schedule-respecting and requires Greedy Coordinate Gradient with a full harmful target), and it generalizes across SFT and preference-optimized (VRPO) models. Three findings emerge. First, the vulnerability is *irreducibly two-component*: re-masking alone (4.4%) and prefix alone (5.7%) both fail. Second, for step-persistent perturbations (a single $L_g \times V$ tensor reused across denoising steps), gradient optimization via a differentiable Gumbel-softmax chain consistently *degrades* ASR (41.5% vs. 76.1%), because continuous perturbations push token distributions off-manifold. Third, A2D, a strong published dLLM defense, is *more* vulnerable to TrajHijack (89.9%) than the undefended model (76.1%): its silent-refusal training removes the contextual resistance that trajectory-level attacks must overcome, an effect we call the **Defense Inversion Effect**.
Chat is not available.
Successful Page Load