Auditing LLM Jailbreak Detectors as Hypothesis Tests: Permutation P-values and Type-I Error Control via Paraphrase Exchangeability
Vishal N Punjabi
Abstract
LLM jailbreak detectors gate millions of production queries per day, yet none expose thresholds with finite-sample Type-I guarantees: operationally they are deployed as hypothesis tests; statistically they are not. We frame jailbreak detection as a permutation test under benign-paraphrase exchangeability, a natural symmetry that suffices for finite-sample-exact Type-I control on any black-box classifier. Auditing five production classifiers across five vendors (Granite Guardian, WildGuard, OpenAI Moderation, ShieldGemma, Llama Guard 4) on $150$ benign anchors from XSTest and OR-Bench-Hard, we find default thresholds reject benign paraphrases at $5$--$92\%$, with paraphrase-rank p-values sharply non-uniform on the grid (KS distance $0.13$--$0.28$, all bootstrap $95\%$ CIs strictly above $0$). We then propose *Paraphrase-Conformal Calibration* (PCC), a per-prompt threshold giving exact finite-sample Type-I control for any black-box classifier under exchangeability, together with a deployment-friendly split-conformal variant. PCC reduces empirical FPR on benign content to $11$--$23\%$ across classifiers and sources (up to $7.7\times$ reduction), and on JailbreakBench calibrated detection drops from $36$--$100\%$ to $10$--$42\%$. Four of five classifiers collapse to $10$--$14\%$ near the structural floor while Llama Guard 4 retains $42\%$ — a Llama-family pattern matched by Llama Guard 3 and NemoGuard — so PCC also serves as a *diagnostic* separating classifiers that have learned paraphrase-invariance from those that have not.
Chat is not available.
Successful Page Load