Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator
Abstract
Randomized experiments (or A/B tests) are widely used to evaluate interventions in dynamic systems such as recommendation platforms, marketplaces, and digital health. In these settings, interventions affect both current and future system states, so estimating the global average treatment effect (GATE) requires accounting for temporal dynamics, which is especially challenging in the presence of nonstationarity; existing approaches suffer from high bias, high variance, or both. In this paper, we address this challenge via the novel Truncated Policy Gradient (TPG) estimator, which replaces instantaneous outcomes with short-horizon outcome trajectories. The estimator admits a policy gradient interpretation: it is a truncation of the first-order approximation to the GATE, yielding provable reductions in bias and variance in nonstationary Markovian settings. We further establish a central limit theorem for the TPG estimator and develop a consistent variance estimator that remains valid under nonstationarity with single-trajectory data. We validate our theory with two real-world case studies. The results show that relative to existing approaches, a well-calibrated TPG estimator can achieve a favorable balance between bias and variance in nonstationary settings, highlighting the value of the policy-gradient perspective for designing effective estimators under complex dynamics.
Lay Summary
A/B tests are widely used to understand whether a new product feature, policy, or intervention works. In many real-world settings, such as recommendation platforms, online marketplaces, and digital health systems, an intervention can affect not only what happens immediately, but also what happens later. For example, changing what a user sees today may influence their future behavior, which in turn changes future outcomes. This makes it difficult to measure the overall effect of an intervention, especially when the system is itself changing over time. This paper proposes a new method, called the Truncated Policy Gradient (TPG) estimator, for measuring treatment effects in such dynamic settings. Instead of looking only at immediate outcomes, the method uses short future outcome paths to capture some of the longer-term impact of an intervention. We show that this approach can provide more reliable estimates by balancing two common problems: estimation bias and variance. We also develop a way to quantify uncertainty from a single experimental dataset. Through two real-world case studies, we show that the proposed method can outperform existing approaches in changing, dynamic environments.