Safety Washing: A Framework for Identifying Superficial Alignment in Large Language Models
Abstract
As large language models (LLMs) are deployed in healthcare, legal services, education, and civic infrastructure, the gap between apparent benchmark safety and real-world behavior has become an urgent concern. We introduce "safety washing," a term analogous to greenwashing in environmental policy, to describe the systematic overestimation of LLM safety by current evaluation infrastructure. We formalize this gap using a probabilistic framework distinguishing structural decoupling from ordinary distribution shift, present a taxonomy of five mechanisms through which safety washing manifests, and synthesize these into the Safety Washing Index (SWI): a multi-criteria audit rubric with four formal design requirements and concrete empirical instantiations. We show that existing published results already provide evidence for each of the five mechanisms, and discuss implications for third-party auditing, model card disclosure, and regulatory standards.