Stress-testing AI incident escalation: design patterns that drive systematic under-detection
Abstract
As AI incident reporting requirements emerge in regulation, a critical question remains unanswered: do the escalation criteria in these frameworks actually detect the incidents they are designed to catch? We develop an eight-criterion escalation framework for AI incidents, then stress-test it against ten real and structured AI incidents spanning cyber, CBRN, psychological harm, and multi-agent risk domains. The stress testing reveals that escalation criteria depend on a three-layer infrastructure: definitional foundations, data availability, and trigger logic. Gaps in upstream layers propagate downstream, producing systematic blind spots. We identify design patterns in current frameworks, including the EU AI Act and California's SB 53, that lead to under-detection: individual-only incident assessment that misses harms emerging from accumulation; and (4)~discrete-event architecture applied to standing conditions, rendering ongoing population-level harms, such as psychological harm from human--AI interaction, invisible to escalation criteria. Drawing on financial services precedent, we propose tolerance-based monitoring as a design principle for these standing conditions.