Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression
Abstract
We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A central challenge in this setting is confounding: treatment assignment often depends on covariates, creating selection bias that makes direct regression of the response on treatment unreliable. To address this issue, we propose a two-stage kernel ridge regression method. In the first stage, we learn a model for the response as a function of both treatment and covariates; in the second stage, we use this model to construct pseudo-outcomes that correct for distribution shift, and then fit a second model to estimate the treatment effect. Although the response varies with both treatment and covariates, the induced effect function obtained by averaging over covariates is typically much simpler, and our estimator adapts to this structure. Our optimal learning bounds are achieved without estimating the conditional treatment density, thereby bypassing a major bottleneck in existing methods. Furthermore, we introduce a fully data-driven model selection procedure that achieves provable adaptivity to both the unknown degree of overlap and the spectral decay of the underlying kernel.
Lay Summary
Many scientific and policy questions ask how outcomes change as the amount of an intervention changes, such as a drug dose, exercise intensity, or funding level. This is hard to learn from observational data, because people who receive different amounts often differ in many other ways. A simple comparison can therefore mistake pre-existing differences for the effect of the treatment. We develop a two-step learning method for estimating this dose-response curve. First, it learns how outcomes depend on both the treatment level and background information about each person. Then it uses that learned relationship to adjust for imbalance in the data and estimate how the average outcome would change across treatment levels. Unlike many existing methods, our approach does not require directly estimating how likely each person was to receive each possible treatment level, which can be unstable when the data are complex or sparse. Nevertheless, our theory shows that the estimator is as accurate as possible in a precise statistical sense under our assumptions, and our numerical experiments support these guarantees.