Distributionally Robust Causal Abstractions
Abstract
Causal Abstraction (CA) theory provides a principled framework for relating causal models that describe the same system at different levels of granularity while ensuring interventional consistency between them. Recent methods for learning CAs, however, assume fixed and well-specified exogenous distributions, leaving them vulnerable to environmental shifts and model misspecification. In this work, we address these limitations by introducing the first class of distributionally robust CAs and their associated learning algorithms. The latter cast robust causal abstraction learning as a constrained min-max optimization problem with Wasserstein ambiguity sets. We provide theoretical guarantees for both empirical and Gaussian environments, enabling principled selection of ambiguity-set radii and establish quantitative guarantees on worst-case abstraction error. Furthermore, we present empirical evidence across different problems and CA learning methods, demonstrating our framework’s robustness not only to environmental shifts but also to structural and intervention mapping misspecification.
Lay Summary
Many real-world systems can be studied at different levels of detail: for example, individual measurements can be grouped into a summary, or low-level image pixels can be mapped to more meaningful features. Causal abstraction studies how to build these simplified views while preserving the causal relationships needed to answer “what if?” questions. Existing methods usually assume that the conditions under which the data were collected are fixed and correctly specified, which can make the learned abstraction unreliable when the environment changes. To tackle this challenge, we introduce DiRoCA, a method for learning causal abstractions that are designed to remain reliable across a range of plausible environments rather than only one observed setting. The key idea is to train the abstraction against worst-case but realistic changes in the hidden sources of variation in the data. Across synthetic, semi-synthetic, and real-world examples, DiRoCA produces abstractions that are more stable under distribution shifts and several kinds of model misspecification. This can make simplified causal models more trustworthy when they are used outside the exact conditions in which they were learned.