Why do Irrelevant Instructions Inhibit Refusal?
Abstract
Large language models (LLMs) are finetuned to produce responses that satisfy multiple criteria, including instruction-following and safety. Cases where models' priorities conflict require prioritization of some objectives over others, and this can be adversarially leveraged. Here, we demonstrate that adding neutral instructions can decrease LLM refusal of harmful goals across model types and scales. In one model (Qwen3-1.7B), we show how a number of linear subspaces generated by differences-in-mean activations related to refusal, harm, and instructions contain overlapping and separate information about harm and instructions across layers. We additionally find evidence that the geometry of refusal is shifted by the presence of neutral instructions. In line with this, steering with a vector generated without instructions present was insufficient to reinstate refusal behavior in a neutral instruction context, but steering with instruction-related vectors unrelated to refusal were able to reinstate refusal behaviors. These results give valuable insights about challenges and possibilities in utilizing linear perturbations in LLMs to counteract safety failure modes.