Refusal Lives Downstream of Persona in Chat Models
Abstract
Linear directions in activation space have been identified for both refusal behavior and persona traits in instruction-tuned chat models, but the two have largely been studied as separate mechanisms. In this work, we find that they are directionally dependent: personas can condition refusal. We identify directions encoding relational persona and show that steering along them exerts coherent, broad-spectrum control over generation. This control extends to refusal: steering toward a "compliant" persona suppresses it via late-layer persona projections. Layer-resolved interventions localize the effect to late layers: refusal is recoverable either by reintroducing it there or by ablating the suppressing persona projection, but not by reintroducing it at an earlier layer. These results suggest that action-level refusal representations interact with late-layer identity-level persona representations. More broadly, they show that safety behavior in chat models is mediated by interactions between interpretable directions, rather than by a single refusal direction alone.