Causal Matrix Completion under Multiple Treatments via Mixed Synthetic Nearest Neighbors
Abstract
Synthetic Nearest Neighbors (SNN) provides a principled solution to causal matrix completion under missing-not-at-random (MNAR) by exploiting local low-rank structure through fully observed anchor submatrices. However, its effectiveness critically relies on sufficient data availability within each treatment level, a condition that often fails in settings with multiple or complex treatments. In this work, we propose Mixed Synthetic Nearest Neighbors (MSNN), a new entry-wise causal identification estimator that integrates information across treatment levels. We show that MSNN retains the finite-sample error bounds and asymptotic normality guarantees of SNN, while enlarging the effective sample size available for estimation. Empirical results on synthetic and real-world datasets illustrate the efficacy of the proposed approach, especially under data-scarce treatment levels.
Lay Summary
Past data often tells us what happened after one action, but not what would have happened under other actions. For example, a platform may ask how a user would react to a different ad exposure, or a policymaker may ask how cigarette sales would change under a different tobacco policy. This is hard when there are many action levels and some levels are rare, because existing methods often need enough examples from the same level. We propose Mixed Synthetic Nearest Neighbors, a method that estimates missing “what-if” outcomes by borrowing information across action levels, even when the observed data is biased. The key idea is that people, products, or states often keep stable underlying traits across actions, so common actions can help learn about rare ones. We show that this greatly increases the amount of usable data while keeping the reliability guarantees of earlier methods. In simulations and a tobacco-policy case study, our method estimates rare-action effects more often and more accurately than the original approach.