A Statistical Framework for Mechanistic Claims in Neural Networks: The Predict--Intervene--Validate Pipeline
Zacharie Bugaud
Abstract
Mechanistic interpretability makes causal claims about neural network internals, yet lacks a standardised statistical framework for validating them. We propose PIV (Predict-Intervene-Validate), a three-stage pipeline for establishing mechanistic claims with quantifiable confidence: (1) Predict: derive a falsifiable prediction from the mechanism and test it blind on held-out models, requiring pre-registration of the certificate before test data generation; (2) Intervene: perform causal interventions (weight surgery, clamping) with proper controls and dose-response analysis, reporting effect sizes and confounds; (3) Validate: test OOD generalisation of the claim itself, not just the model, using hyperparameter-frozen transfer. We instantiate PIV on a concrete mechanism (input-invariant dimensions in Elman RNNs) and demonstrate its discriminative power: blind prediction achieves AUROC $0.97$ on $300$ fresh models, with the pre-registered precision-optimal operating point yielding $97.8\%$ precision and $92.0\%$ recall, while $13$ of $15$ candidate hypotheses are rejected against pre-registered thresholds; confound-controlled surgery produces $393/393$ causal breaks on the hypothesised mechanism ($p<10^{-42}$); and hyperparameter-frozen OOD transfer at $12.5\times$ scaling achieves $160/160$ ($p<10^{-71}$, beta-binomial). We derive sample-size guidelines for each PIV stage via power analysis, enabling practitioners to design adequately powered studies. The framework applies to any population-level mechanistic claim and is released as a checklist for mechanistic interpretability submissions, together with the piv Python package and a hosted demo notebook.
Chat is not available.
Successful Page Load