The Pareto-optimal Trade-off between Regret and Statistical Inference in Linear Stochastic Bandits under Safety Constraints
Abstract
Lay Summary
Many automated systems learn by trial and error: they try an option, see the result, and adjust. A medical trial assigning treatments to patients works this way. Such systems usually chase one goal: giving each patient the best treatment we can along the way. But two others matter just as much. First, we want to understand why a treatment works, so the conclusions hold up for future patients. Second, every choice must stay safe, never exposing anyone to unacceptable risk. These goals pull against each other: gathering the data needed for solid understanding means occasionally trying options that look less appealing or riskier. We asked how good a system can possibly be at all three at once, and whether any method reaches that limit. We proved the exact boundary of what is achievable and designed a method that provably hits it, while keeping risk low enough to be effectively negligible. This gives designers of high-stakes learning systems a principled way to balance performance, trustworthy conclusions, and safety, instead of trading them off by guesswork.