Behavior-Only Deployment Certification Requires Bridge Assumptions
Shuo L Liu
Abstract
A model can pass every behavioral evaluation an auditor runs and still fail a deployment safety claim, not because the auditor made a statistical mistake, but because the evaluation never touched the histories on which the claim depends. We formalize this failure mode for agentic systems. A fixed evaluator $Q$ observes transcripts on an evaluation-induced support and asks whether a deployment-risk property $R_{\rm deploy}(\pi)\le r$ holds. If deployment contains a bad set $B$ outside that support and the admissible policy class permits off-support variation, a safe policy and an unsafe policy can induce the same transcript law under $Q$, with the unsafe policy copying the safe policy on every evaluator-reachable history and failing on $B$. If $B$ is merely rare rather than unreachable, constant-power behavior-only detection requires $\Omega(1/\varepsilon)$ episodes when the evaluator reaches $B$ with probability $\varepsilon$. Behavioral tests, protocol-visible logs, red-team transcripts, and preference scores remain useful for debugging and distribution-scoped risk estimation. They become deployment certificates only when support expansion, structural restrictions, trusted non-behavioral evidence, or other bridge assumptions link evaluation histories to deployment histories.
Chat is not available.
Successful Page Load