Partial Identification under High-Dimensional Potential Outcomes and Confounders via Optimal Transport
Abstract
Partial identification provides informative causal guarantees when point identification is impossible, but existing approaches based on optimal transport (OT) become computationally and statistically intractable in high-dimensional settings. This limitation is particularly severe when both potential outcomes and confounders are high-dimensional, where classical OT-based bounds suffer from the curse of dimensionality and unfavorable convergence rates. To address this challenge, we propose a novel estimator that decomposes the transport problem into a low-dimensional signal subspace and a high-dimensional residual subspace. Unlike existing projection-based methods that discard residual information, we recover the residual transport energy using the Sliced Wasserstein distance, which is computationally efficient and robust to high dimensions. We establish interpretable conditions controlling the approximation gap based on residual structure and provide a data-driven rule for signal dimension selection. Empirical results show that our estimator consistently outperforms projection-only baselines by recovering lost transport energy, yielding more informative causal bounds while remaining computationally tractable in high dimensions.
Lay Summary
In many studies, such as medical studies, we want to know what would have happened to the same person under different choices. The data only show one outcome for each person, so the effect often cannot be pinned down exactly. Instead of pretending there is one exact answer, partial identification gives a range of answers that are consistent with the data. Existing ways to compute these ranges can use optimal transport, a mathematical way to compare two distributions, but they become unreliable when each person is described by many variables. We propose a more stable method that first finds the few directions where the two groups differ most, and then uses many simple one-dimensional comparisons to account for the remaining differences. This keeps the computation manageable while avoiding the loss of information caused by using only a low-dimensional projection. We prove when this calculation is valid and when it is close to the full comparison. In simulations and a real right-heart catheterization dataset, the method gives more informative bounds than projection-only alternatives. This can help researchers build causal bounds in high-dimensional studies without relying on overly strong modeling assumptions.