Budgeted Active Experimentation for Treatment Effect Estimation from Observational and Randomized Data
Abstract
Estimating heterogeneous treatment effects is central to data-driven decision-making, yet industrial applications often face a fundamental tension between limited randomized controlled trial (RCT) budgets and abundant but biased observational data (OBS) collected under historical targeting policies. Although observational logs offer the advantage of scale, they may suffer from severe policy-induced imbalance and overlap violations, rendering standalone estimation unreliable. We propose a \textit{budgeted active experimentation} framework that iteratively collects informative randomized samples for causal effect estimation via active sampling. By leveraging observational signals, we develop an acquisition function targeting uplift estimation uncertainty, domain discrepancy, and overlap deficits to select the most informative units for randomized experiments. We establish finite-sample deviation bounds, asymptotic normality via martingale CLTs, and minimax lower bounds showing near-optimality in the linear representation setting. Experiments on synthetic datasets support our theoretical findings, and further extensions to industrial neural network-based uplift modeling scenarios show that active sampling can improve sample efficiency over random sampling under limited RCT budgets.
Lay Summary
Many decision-making systems need to estimate who is likely to benefit from an intervention, such as a discount, a recommendation, or a service upgrade. Randomized experiments provide reliable evidence for this question, but they are often costly and can only include a limited number of users. Historical data are much more abundant, but they may be biased because past actions were chosen by existing business rules rather than by random assignment. This paper studies how to combine these two sources of information. We do not use biased historical data as direct causal evidence. Instead, we use it to decide where new randomized experiments should be run. The proposed method selects users whose experimental outcomes are expected to be most informative, especially in parts of the population where historical data provide weak or unreliable evidence. By directing a limited experimental budget toward more useful samples, the method learns treatment-effect models more efficiently. Experiments on synthetic data and large-scale industrial data show that this strategy can achieve better estimation performance than standard random sampling under the same budget.