Causal-EPIG: Causally Aligned Active CATE Estimation
Abstract
Estimating the Conditional Average Treatment Effect (CATE) is constrained by the high cost of obtaining outcome measurements, making active learning valuable. However, conventional strategies suffer from a fundamental objective mismatch: they reduce uncertainty in model parameters or observable outcomes rather than the unobservable causal quantities of interest. We address this via the principle of causal objective alignment, positing that acquisition functions should target potential outcomes or CATE directly. We operationalize this through Causal-EPIG, a framework adapting Expected Predictive Information Gain to quantify uncertainty reduction in causal quantities. We derive two distinct strategies: a comprehensive approach that targets the joint potential-outcome structure, and a focused approach that directly targets the CATE estimand for sample efficiency. We provide theoretical justification for our framework, establishing a formal link between CATE estimation error and posterior uncertainty in causal quantities. Extensive experiments demonstrate that our strategies improve sample efficiency over standard baselines, and crucially, reveal that the preferred strategy is context-dependent, contingent on the base estimator and treatment-effect structure. Our framework thus provides a principled guide for sample-efficient CATE estimation in practice.
Lay Summary
Estimating how treatments affect different people is important in medicine, economics, and public policy, but measuring outcomes can be expensive. When only a limited number of outcomes can be collected, we need to choose the most useful ones. This paper proposes Causal-EPIG, a method for selecting outcome measurements by focusing directly on the causal question: how treatment effects vary across people. Instead of reducing uncertainty about model details or ordinary predictions, it targets the treatment-related quantities we ultimately care about. Experiments show that this can make data collection more efficient, while the best strategy depends on the model and data setting.