High-Dimensional Sensitivity Analysis for Genomic Studies: An Adversarial Framework for Learning Worst-Case Latent Confounders
Abstract
High-dimensional genomics studies are frequently confounded by unmeasured biological processes that obscure disease-specific signals. While existing workflows can estimate these latent confounders, they fail to quantify how robust a discovery is to varying levels of hypothetical confounding. We introduce sensGAN, a deep-learning adversarial framework that systematically explores the confounding spectrum by learning "worst-case" latent variables that nullify the most gene associations under novel predictive-gain constraints. By identifying the minimum confounding strength required to explain away an observed effect, our method shifts the paradigm toward a formal, quantitative sensitivity analysis. In diverse simulations, sensGAN accurately recovers latent structures and outperforms existing methods in identifying confounder-sensitive genes. Applied to human Alzheimer's disease microglia, our framework prioritizes robust disease pathways while successfully isolating signals driven by unmeasured co-occurring neurodegenerative pathologies.
Lay Summary
Genomic studies often identify disease-associated genes that may actually reflect hidden biological processes that researchers did not measure. Existing methods try to adjust for these hidden effects, but they cannot measure how strongly they influence a discovery. We developed sensGAN, a machine learning framework that generates hypothetical “worst-case” hidden factors and tests whether gene discoveries remain significant after adjustment. The method systematically measures how much hidden influence is needed to explain away each result. Applied to Alzheimer’s disease brain data, sensGAN separated robust disease-related genes from signals likely driven by other co-occurring neurodegenerative diseases. This can help researchers prioritize more reliable biological discoveries and improve confidence in genomic studies.