A model of early warning systems in an AI race
Abstract
Tech-race dynamics could delay action on catastrophic risks from future AI systems. An underappreciated reason for this delay is that tech-race dynamics can make warnings of AI risk appear manipulative. Evidence of an imminent AI risk will be insufficient at motivating collective action if rivals suspect the evidence inflates the risk. I model two competing AI labs, each with a private early warning system, deciding whether to undertake a costly safety project and whether to disclose their warning evidence to the other. Two results follow. First, a hysteresis trap in which labs hold risk thresholds too high to act on for fear of losing their position in the race. This result means that even compelling warnings fail to trigger coordinated safety action. Second, selective disclosure aimed at amplifying the severity of a warning can undermine trust in shared evidence. To maximise the credibility of a warning they receive, rivals must be close to verifying evidence for themselves. I also use the model to assess the strategy proofness of current early warning systems and provide a simplified procedure for scoring early warnings based on a published safety case for AI scheming.