The Three-Axis Audit Problem in Iterative LLM Safety Hardening
Abstract
Regulatory frameworks mandate iterative safety evaluation for frontier AI systems but leave the operational details unspecified. We study iterative adversarial supervised fine-tuning (SFT) under three levels of certification scrutiny. First, iterative co-evolution converges: a simultaneous attacker-defender loop reduces attack success rate (ASR) to safe levels within three rounds under both SFT and multi-turn group relative policy optimization (GRPO) attackers, replicating prior iterative co-evolutionary results. Second, the resulting defense does not generalize across attack families. Defenders hardened against PAIR, Crescendo, or AutoDAN in isolation retain double-digit held-out ASRs against non-training families. Third, even under multi-family co-evolution, the benign-mix ratio dissociates three audit axes (direct safety, benign refusal, and cross-family robustness). No ratio we tested clears all three, and two of three possible two-axis audits silently certify non-robust checkpoints. We propose two governance instruments: a three-axis joint reporting protocol, and a hardened_against Model Card field declaring the attack families a checkpoint was certified against. Both align with EU AI Act Articles 9 and 13.