Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute
Abstract
Frontier AI labs increasingly monitor chat interactions for safety-relevant harms and to gather information on broad usage patterns, but the same infrastructure could be repurposed for surveillance, and monitored parties have no technical means to verify that a stated monitoring scope matches the actual one. We propose a two-party protocol in which a monitor and a monitored party jointly sign a Plan that is executed by an open-source LLM classifier inside an attested Trusted Execution Environment (TEE), so that out-of-scope queries are never evaluated and scope changes require fresh signatures from both parties, which we call Verifiably-Scoped Monitoring. As a primary application, the protocol reconciles safety monitoring with zero-data retention (ZDR): the company can verify that the lab's monitoring stays within the jointly-signed scope, while the lab still receives the monitoring outputs it needs.