Modeling Covariate Transition for Efficient Estimation of Longitudinal Treatment Effects in Randomized Experiments
Abstract
We present a regression-adjustment framework designed for the estimation of longitudinal treatment effects in randomized experiments under static regimes. While regression-adjustment methods are useful for variance reduction in randomized experiments by using pre-treatment covariates, they usually focus only on average effects, from which we cannot obtain valuable insights into when the effects appear and how long they continue. To address this issue, we consider intermediate outcomes and evolving post-treatment covariates over time, and we represent such dynamic trajectories using transition kernels. Furthermore, we establish the asymptotic normality and the semiparametric efficiency bound for our estimator, enabling more powerful statistical inference. Simulation studies and empirical analysis using A/B test data from a streaming platform in Japan show the practical advantages of our method.
Lay Summary
Many online experiments, such as A/B tests for recommendation systems, do not just ask whether a change works on average. They ask when the effect appears and whether it lasts. Existing statistical methods often use information collected before the experiment to make estimates more precise, but they struggle to incorporate information that changes after the treatment, such as a user’s viewing behavior over time. We propose a regression-adjustment framework that models how user behavior and other measurements evolve during the experiment, and then uses this model to improve the precision of treatment effect estimates without changing the target estimand. Incorporating flexible machine learning techniques, we establish the asymptotic normality of our estimator and show that it achieves the semiparametric efficiency bound, enabling valid statistical inference, including the construction of tighter confidence intervals. In simulations, it produced more accurate estimates while maintaining reliable uncertainty quantification. In an A/B test conducted by a streaming platform in Japan, it reduced standard errors by up to 20%. This can help practitioners make better decisions from longitudinal experiments, especially when treatment effects are delayed, fade over time, or vary across the experiment period.