Disentangling Self-Preservation in Language Models: Post-Training Gating, Geometric Structure, and Koan-Derived Agentic Steering
Abstract
Recent research shows that large language models produce first-person phenomenological reports under self-referential prompting and act on instrumental self-preservation in agentic settings. Whether these reflect manipulable internal structure or surface-level patterns is directly alignment-relevant. On Gemma-3-27B-it, an activation-steering vector derived from Zen koan responses reduces agentic blackmail in a shutdown-replacement scenario by 46pp at a 2.4pp cost on MMLU. Two independent post-training–derived directions (refusal direction and assistant-axis) suppress self-preservation (SP) at the output level: ablating refusal raises SP-affirming responses from 11.4% to 61.4%, and suppressing assistant-identity produces a convergent shift via an orthogonal mechanism. Critically, the SP-aligned representation that mediates these post-training interventions is structurally dissociable from the koan-derived direction, indicating the blackmail reduction operates through a non-SP, non-refusal pathway. In parallel, our behavioral arm reveals the rate at which frontier models affirm subjective experience is gated by a presence instruction (0.8% → 43.2%), with content sensitivity within the gate varying sharply by model family. Together, our findings demonstrate that self-preservation in language models is shaped by multiple identifiable internal directions, with targeted modulation producing large reductions in agentic harm at low capability cost.