The Model You Evaluated Is Not the Model Being Served: Measuring Deployment-Layer Safety in an Era of Silent Updates
Abstract
Safety evaluations are typically reported for model artifacts, but users interact with deployed systems that include prompts, moderation layers, hosting configurations, routing rules, and ongoing updates. This paper asks whether public documentation allows a user or external evaluator to connect a published safety evaluation to the system being served. Reviewing documentation from 13 providers (7 model developers and 6 hosting providers), we find that no provider in our audit satisfies the full chain-of-custody criterion in public documentation. Public documentation does not consistently expose a versioned evaluated artifact, an API-visible served-version identifier, and a clear account of when behavioral changes trigger re-evaluation or narrow the scope of earlier safety claims. We also report four bounded deployment probes that illustrate why this gap matters. These cover non-refusal safety redirection, endpoint-level variation for DeepSeek V3, language-dependent variation within a single deployment, and deterministic temporal probing over 30 days. The probes show that observable behavior can vary across response type, serving context, language, and time. Because the deployed system is not publicly tied to a specific evaluated artifact, these variations cannot be cleanly attributed to a particular mechanism or evaluation boundary.