Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing
Abstract
Beyond conditional average treatment effects, treatments may impact the entire outcome distribution in covariate-dependent ways, for example, by altering the variance or tail risks for specific subpopulations. We propose a novel estimand to capture such conditional distributional treatment effects, and develop a doubly robust estimator that is minimax optimal in the local asymptotic sense. Using this, we develop a test for the global homogeneity of conditional potential outcome distributions that accommodates discrepancies beyond the maximum mean discrepancy (MMD), has provably valid type 1 error, and is consistent against fixed alternatives---the first test, to our knowledge, with such guarantees in this setting. We then provide a test that aggregates evidence across a grid of kernel-bandwidth choices. Furthermore, we derive exact closed-form expressions for two natural discrepancies (including the MMD), and provide a computationally efficient, permutation-free algorithm for our test.
Lay Summary
Many studies ask whether a policy or treatment changes the average outcome, but averages can hide important effects. For example, two medical treatments may look equally good on average, even if one greatly increases the chance of severe outcomes for certain patient groups. We focus on such conditional distributional treatment effects: "conditional" because a treatment's effect may depend on an individual's observed background---such as demographics, medical history, or biomarkers, and "distributional" because a treatment may affect any feature of the possible outcomes, including the average, variability, or the chance of rare severe complications. We develop an efficient method for estimating such effects by comparing treated and untreated groups, adjusting for the observed differences between them. Our method can use flexible machine learning algorithms, and features a safety net known as "double robustness": it remains reliable if even one of the two main underlying data patterns is modeled correctly. We use this to develop a principled test for whether any subgroup of the population experiences a distributional treatment effect, and mathematically prove that, under reasonable conditions, our test avoids false alarms and gets better at detecting real effects as dataset sizes grow. We also provide computationally efficient algorithms for this test, making our approach practical for large datasets. In experiments, our approach detects effects that average-based analyses can miss and helps show where in the population those effects occur.