Robust Sequential Experimental Design for A/B Testing
Abstract
Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.
Lay Summary
A/B testing is widely used by technology companies to compare different versions of a product, such as a webpage, recommendation system, or app feature. A key challenge is how to assign users to different versions efficiently so that reliable conclusions can be drawn with fewer samples. Existing experimental design methods can improve efficiency, but they often depend on statistical models being correctly specified. In practice, these models may be imperfect, which can lead to inaccurate conclusions. This paper develops a robust sequential experimental design method for A/B testing that remains reliable even when the working model is misspecified. The proposed framework can handle both one-step decision problems, such as contextual bandits, and multi-step decision problems, where decisions and outcomes evolve over time. We show theoretically that the method controls the worst-case error in estimating treatment effects. We also demonstrate through simulations and real-world data from a major technology company that the method improves the reliability and efficiency of A/B testing.