What AI Governance Needs from Mechanistic Auditing
Abstract
As AI systems advise on medical diagnoses, legal cases, and financial decisions, regulators demand verifiable safety guarantees. Behavioral testing alone cannot provide these guarantees—we need to understand model internals. But what does 'understanding' mean for governance purposes? This paper argues that mechanistic auditing must satisfy eight requirements: domain experts define safety properties, technical auditors localize and validate relevant circuits, explanations are translated with coverage estimates, and corrections are verified for persistence and side effects. Current methods fail on two critical requirements—automated circuit localization at scale and faithful translation to domain concepts with quantified coverage—making these the field's central open problems. We discuss applications in verifying reasoning faithfulness, predicting capa- bility emergence, and mapping dangerous compo- sitions, while analyzing failure modes including interpretability illusions and adversarial evasion.