Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
Abstract
Governance frameworks increasingly rely on documented evaluation evidence, but benchmark-validity audits can themselves be fragile measurement pipelines. This paper identifies five classes of pipeline-level failures that can silently alter or reverse audit conclusions, including silent no-op perturbations, regex-extraction artifacts, non-faithful scoring, broken bootstrap pairing, and metric-archetype mismatch. We demonstrate these failures through a self-audit of five safety benchmarks and two open-weight instruction-tuned models, showing that all ten benchmark-model cells fall into non-confirmatory buckets under a unified six-point due-diligence gate. Rather than proposing a new leaderboard metric, we position the gate as a withholding and disclosure protocol for assurance-grade AI governance evidence: benchmark-audit results should only be treated as confirmatory when the scorer path, uncertainty estimate, benchmark archetype, and repair-regression checks are explicitly validated and disclosed.