Sequential Hypothesis Testing for Modern AI Systems: Computation, Robustness, and Online Monitoring
Abstract
Modern AI systems are increasingly deployed as interactive, evolving, and high-dimensional systems rather than static predictors. Their behavior may change after deployment because of shifts in user populations, prompts, retrieval sources, tools, feedback loops, or model updates. This creates sequential hypothesis testing problems: one must detect changes quickly while controlling false alarms under repeated monitoring, limited memory, and limited computation. In this talk, I will discuss this problem through the lens of sequential change-point detection. I will start with classical likelihood-ratio based methods, then explain the tradeoff among statistical power, model robustness, computation, and memory. I will next discuss adaptive and online learning methods, followed by modern nonparametric testing ideas based on MMD, kernel witnesses, including Stein critics, and online kernel CUSUM. The talk will close by connecting these tools to post-deployment monitoring of modern AI systems, including large language model systems.