Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression
Abstract
We address the problem of causal effect estimation in the presence of hidden confounders using nonparametric instrumental variable (IV) regression. An established approach is to use estimators based on learned \emph{spectral features}, that is, features spanning the top singular subspaces of the operator linking treatments to instruments. While powerful, such features are agnostic to the outcome variable. Consequently, the method can fail when the true causal function is poorly represented by these dominant singular functions. To mitigate, we introduce Augmented Spectral Feature Learning, a framework that makes the feature learning process outcome-aware. Our method learns features by minimizing a novel contrastive loss derived from an augmented operator that incorporates information from the outcome. By learning these task-specific features, our approach remains effective even under spectral misalignment. We provide a theoretical analysis of this framework and validate our approach on challenging benchmarks.
Lay Summary
Many scientific and policy questions ask what would happen if we changed something, such as whether more education raises earnings. This is hard when hidden factors, like family background or ability, affect both the choice being studied and the outcome. Instrumental-variable methods address this problem by using additional observed variables that help separate cause from coincidence. In modern applications, however, the relationships between all these variables can be highly nonlinear and complicated, so it is important for machine learning methods to learn useful summaries of the data. Existing methods often learn summaries that capture the strongest relationships between variables, but these may not be the relationships that matter for estimating the causal effect. Our work changes this learning step so that the method also uses information from the outcome when deciding which summaries to learn. In effect, it searches for information that is both statistically reliable and useful for the causal question. We prove when this helps and test it on several challenging examples, including image-based data and policy-evaluation tasks for decision-making systems. The result is a more reliable way to estimate cause-and-effect relationships in difficult settings where standard methods can focus on the wrong patterns.