Design-Based Anytime-Valid Inference for Randomized Experiments with Delayed Outcomes and Staggered Entry
Abstract
Lay Summary
Real-world A/B tests are often more complicated than they first appear. A marketing intervention, for example, might make some customers purchase sooner while reducing the fraction who eventually purchase at all. At short time scales the treatment may therefore look beneficial, while at longer time scales it may be harmful. Revenue adds another layer: a treatment might increase the number of purchases and accelerate when they occur, but steer customers toward cheaper products and reduce total revenue. This creates a challenge for continuous monitoring. Although the eventual outcome is often the natural target, estimating it during an ongoing experiment requires assumptions that may be hard to justify. In practice, units enter experiments over time rather than all at once, and users who enter at different times may differ substantially. They may also experience different product versions, market conditions, or external environments. We therefore ask what can be monitored continuously without relying on such assumptions. Our answer is the cumulative reward, such as total revenue, that would have been observed by each calendar time had all entered units been assigned to treatment versus control. We construct confidence sequences, uncertainty bands that remain valid under continuous monitoring, for this treatment-control difference. The resulting quantity is directly interpretable: it compares the total reward generated under treatment and control as the experiment unfolds, without requiring a probability model for a hypothetical population.