Addressing Instrument-Outcome Confounding in Mendelian Randomization through Representation Learning
Abstract
Mendelian Randomization (MR) is a prominent observational epidemiological research method, designed to address unobserved confounding when estimating causal effects. It is closely related to instrumental variable (IV) methods, where genetic variants serve as instruments to infer causal relationships from observational data. However, the core assumptions required for valid IV analysis---particularly the independence between instruments and unobserved confounders---are untestable and often violated in practice. In MR, such violations commonly arise when genetic variants are correlated with environmental factors (e.g., population stratification and assortive mating), leading to confounding between instruments and outcomes. At the same time, MR studies increasingly include data collected across multiple environments or populations, providing an opportunity to address these violations. Leveraging this setting, we propose a representation learning framework that exploits multi-environment data to recover latent exogenous components of genetic instruments suitable for causal inference. We provide theoretical insights into when and how the learned components can act as valid instruments, and we demonstrate the effectiveness of our approach through simulations and semi-synthetic experiments using genetic data from the All of Us Biobank.
Lay Summary
Epidemiologists often need to know whether a risk factor causes a disease, for example, whether being overweight causes type 2 diabetes. When randomized clinical trials are impossible or unethical, researchers turn to Mendelian Randomization, a technique that treats the genes we randomly inherit as "natural experiments" to tease apart cause from coincidence. The catch: which genes people carry isn't completely random across populations --- ancestry and environment shape genetic patterns in ways that can quietly distort the conclusions. We show that contrasting genetic data across multiple populations lets us mathematically separate the truly random part of a genetic variant from the part shaped by population history. We propose a machine-learning model to do this separation automatically, and show that the model can produce a "cleaned" genetic signal that is suitable for estimating cause and effect. Through simulations, we validate that our method gives more accurate answers than existing tools, and we demonstrate it on real genetic data from All of Us, a biobank with participants from many backgrounds in the United States. As large and diverse biobanks grow, our approach helps researchers learn about the causes of disease from them without being misled by population differences.