When Can We Conclude a Dangerous Capability is Absent? Statistical Foundations for AI Capability Evaluation
Kaustubh Bukkapatnam ⋅ Siddharth Karuturi
Abstract
Regulatory frameworks—including the EU AI Act, the US Executive Order on AI, and voluntary safety commitments from leading developers—increasingly mandate pre-deployment evaluation of AI systems for dangerous capabilities. Current evaluation practice generates positive evidence readily (demonstrating that a model can do something), but rarely provides statistically rigorous negative evidence (demonstrating that a model cannot do something at policy-relevant rates). We formalise the problem of negative capability claims as a one-sided hypothesis test: given a safety threshold $p_0$ and $n$ independent evaluation trials with $k$ successes, under what conditions can we rigorously conclude that the model’s true capability rate $p < p_0$? We derive (i) exact finite-sample tests and power curves, (ii) a sequential testing procedure based on the sequential probability ratio test (SPRT) that achieves the same guarantees with up to 40\% fewer trials on average when the capability is truly absent, and (iii) Bonferroni and Simes corrections for evaluating $m$ capabilities simultaneously. We apply these tools to published evaluation data and show that several widely-cited evaluations reporting negative capability findings are statistically underpowered at the thresholds implied by their accompanying policy claims.
Chat is not available.
Successful Page Load