Federated Causal Inference on Multi-Site Observational Data via Propensity Score Aggregation
Abstract
Causal inference typically assumes centralized access to individual-level data. Yet, in practice, data are often decentralized across multiple sites, making centralization infeasible due to privacy, logistical, or legal constraints. We address this problem by estimating the Average Treatment Effect (ATE) from decentralized observational data via a Federated Learning (FL) approach, allowing inference through the exchange of aggregate statistics rather than individual-level data. We propose a novel method to estimate propensity scores via a federated weighted average of local scores using Membership Weights (MW), defined as probabilities of site membership conditional on covariates. MW can be flexibly estimated with parametric or non-parametric classification models using standard FL algorithms. The resulting propensity scores are used to construct Federated Inverse Propensity Weighting (Fed-IPW) and Augmented IPW (Fed-AIPW) estimators. In contrast to meta-analysis methods, which fail when any site violates positivity, our approach exploits heterogeneity in treatment assignment across sites to improve overlap. We show that Fed-IPW and Fed-AIPW perform well under site-level heterogeneity in sample sizes, treatment mechanisms, and covariate distributions. Theoretical analysis and experiments on simulated and real-world data demonstrate clear advantages over meta-analysis and related approaches.
Lay Summary
Most methods for estimating causal effects assume that all individual-level data can be stored and analyzed in one central place. In many real-world settings, this is not possible because data are held by different hospitals, institutions, or study sites and cannot be shared directly for privacy, legal, or practical reasons. This work develops a federated approach for estimating the average effect of a treatment using data spread across multiple sites. Instead of pooling individual records, sites only share aggregate information. The method combines information from different sites to estimate each person’s probability of receiving treatment, while accounting for differences between sites. The key idea of our approach to use “membership weights,” which estimate how likely a person with certain characteristics is to belong to each site. These weights allow the method to combine site-specific treatment models into a single federated propensity score. That score is then used in standard causal estimators adapted to the federated setting. The approach improves on traditional meta-analysis, where each site estimates its own treatment effect and the results are combined afterward. Meta-analysis can break down when some sites have poor treatment overlap, for example if nearly everyone at one site receives the treatment. The proposed method can borrow information across sites and use differences in treatment practices to improve overlap. The proposed method enables causal effect estimation when individual-level data cannot be centralized. Simulations and real-world analyses show that it performs better than standard meta-analysis and related methods, especially when sites differ in size, patient populations, and treatment assignment practices.