Securing deep third-party frontier AI evaluations
Abstract
Third-party evaluators increasingly need access to a model's internal computations to assess safety properties that black-box testing cannot reliably detect, such as suppressed capabilities, covert mechanisms, and inadequately removed hazardous knowledge. Model providers will not grant such access because it risks exposing trade secrets. This impasse has stalled the development of credible AI oversight. We propose a de facto glass-box evaluation architecture that resolves this tension by combining hardware-rooted confidential computing, a standardised Evaluation Instrumentation Interface that exposes analytical signals without revealing raw parameters, and a governed egress pipeline that bounds extraction risk. We present a formal threat model, analyse residual risks, and describe three deployment topologies that offer different trust distributions to maximise provider participation. Our analysis indicates that the underlying confidential-computing substrate is production-ready; the evaluation-specific instrumentation layer is a concrete design that requires implementation and empirical validation. A scoped pilot facility could materially advance the state of third-party AI evaluation for safety-critical claim classes.