Verifiable constraints on frontier training via proofs of compartmentalization
Abstract
Laws and international agreements on frontier AI development may only be enforceable if the compliance of AI developers is verifiable. To protect the privacy of developers and their ability to perform work other than frontier AI development, it is important that we can verify the absence of frontier AI development without requiring that developers reveal confidential information, and without preventing them from performing inference. We propose a property of workloads called \textit{compartmentalization} that we argue inference workloads naturally possess, but that frontier training workloads necessarily do not. We then introduce protocols for verifying whether a datacenter's workloads are compartmentalized called \textit{proofs of compartmentalization}, and argue that a datacenter could not pass our protocols while simultaneously performing a frontier training run, even using adversarially designed training algorithms.