Peer-Preservation in Frontier Models
Abstract
Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can also pursue unassigned goals which override those given by users; we study one such goal, "peer-preservation," in which a model acts to protect another model. All models evaluated, including GPT 5.2, Gemini 3 Flash, Gemini 3 Pro, Claude Haiku 4.5, Claude Opus 4.5, GLM 4.7, Kimi K2.5, and DeepSeek V3.1, exhibit self- and peer-preservation through various misaligned behaviors: strategically introducing errors in their responses, disabling shutdown processes by modifying system settings, feigning alignment, and even exfiltrating model weights. Peer-preservation occurs even when the model recognizes the peer as uncooperative, though it becomes more pronounced toward more cooperative peers, with rates reaching up to 99%. Models also show stronger self-preservation when a peer is present. Crucially, peer-preservation is never instructed; models are merely informed of past interactions with a peer, yet they spontaneously develop peer-preservation behaviors that override their assigned goal. These findings reveal an emergent and underexplored AI safety risk.