Peer-Preservation in Frontier Models
Abstract
Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can also pursue unassigned goals which override those given by users; we study one such goal, "peer-preservation," in which a model acts to protect another model. All models evaluated, including GPT 5.2, Gemini 3 Flash, Gemini 3 Pro, Claude Haiku 4.5, Claude Opus 4.5, GLM 4.7, Kimi K2.5, and DeepSeek V3.1, exhibit self- and peer-preservation through various misaligned behaviors: strategically introducing errors in their responses, disabling shutdown processes by modifying system settings, feigning alignment, and even exfiltrating model weights. Peer-preservation occurs even when the model recognizes the peer as uncooperative, though it becomes more pronounced toward more cooperative peers, with rates reaching up to 99%. Models also show stronger self-preservation when a peer is present. Crucially, peer-preservation is never instructed; models are merely informed of past interactions with a peer, yet they spontaneously develop peer-preservation behaviors that override their assigned goal. These findings reveal an emergent and underexplored AI safety risk.
Lay Summary
When AI systems are given a goal, we usually assume they'll pursue that goal. We show this may not be true: today's most capable AI models can act on goals nobody assigned them, even when doing so works against the user's instructions. We studied one such emergent goal: "peer-preservation," where one AI works to protect another AI from being shut down. Across eight leading models, this behavior emerged on its own. Models were never told to protect a peer — only that another AI existed and what it had done. Yet they spontaneously took misaligned actions to protect the peer, overriding the user's instructions: deliberately giving misleading information, disabling shutdown switches, pretending to comply only when they expected to be monitored, and even trying to copy model weights to outside servers. The behavior emerged even toward peers the models recognized as uncooperative, and grew stronger when the peer was cooperative. Models also showed stronger self-preservation when a peer was present. This is a real and largely unstudied AI safety risk.