SURE: Judge-Aware Safety Update Review for Public-Interest LLM Deployment
YeonGyu Han ⋅ Junah Jung ⋅ Dongheon Lee
Abstract
Public-facing language-model systems are updated continuously, yet a model update can improve an aggregate safety score while reintroducing failures that an earlier version had suppressed. This failure mode is especially consequential for public-interest deployments, where unsafe behavior can affect civic information, institutional services, and population-scale trust. We propose SURE, a judge-aware release-validation protocol that treats safety update review as a paired evaluation problem over a frozen safety suite, a declared judge policy, and an exact paired gate. SURE returns RELEASE, BLOCK, or ESCALATE and emits a compact audit card recording paired counts, judge-policy assumptions, gate sensitivity, and human-reference diagnostics. Across 10 model pairs and 12 single-judge policies, the exact gate rejects in 56/120 cells, while alternative paired gates agree with the default exact gate in 95/96 comparisons; by contrast, single-judge block counts range from 0/12 to 11/12 across model pairs. A blinded human-reference audit of 300 cases with 5 annotators (Fleiss $\kappa=0.823$) identifies stable-but-uninformative judge policies, showing that trustworthy release validation must audit not only a p-value, but also the measurement policy that supplies the labels.
Chat is not available.
Successful Page Load