Sinkhorn Treatment Effects: A Causal Optimal Transport Measure
Abstract
We introduce the Sinkhorn treatment effect, an entropic optimal transport measure of divergence between counterfactual outcome distributions. Unlike classical quantities such as the average treatment effect, it captures differences across entire distributions. We show that this estimand can be written as a smooth transformation of counterfactual mean embeddings with an appropriate kernel. This characterization allows us to establish first-order pathwise differentiability in general, and second-order pathwise differentiability under the null hypothesis of equal counterfactual distributions. Leveraging this smoothness, we construct debiased estimators and asymptotically valid tests for distributional treatment effects at a fixed entropic regularization parameter. Because the power of the test depends on this unknown parameter, we propose an aggregated test that combines evidence across a grid of regularization choices. Experiments on simulated and image data demonstrate the practical advantages of our estimator and testing procedure.
Lay Summary
Many studies ask whether a treatment "works" by comparing only average outcomes, but averages can hide important changes: for example, a job-training program may leave average earnings unchanged while increasing earnings for some workers and decreasing them for others. These effects appear in the full shape of the outcome distribution, not just its center. We propose a new way to compare the full outcome distributions that would arise under treatment and control. Our method, called the Sinkhorn Treatment Effect (STE), measures how much work is needed to transform the outcome distribution under treatment into the one under control, while accounting for pre-existing differences between treated and untreated groups. We also develop statistical tools, including bias-corrected estimators of STE and hypothesis tests, to detect these distributional treatment effects. Our estimator helps researchers move beyond the question “Did the average outcome change?” toward the richer question “How did the treatment change the whole population?” That can lead to more informative decisions in applications where averages alone miss clinically, socially, or scientifically meaningful effects.