Learning Adaptive Perturbation-Conditioned Contexts for Robust Transcriptional Response Prediction
Abstract
Predicting high-dimensional transcriptional responses to genetic perturbations is challenging because signals are sparse and experimental noise is severe. Existing methods often suffer from mean collapse, achieving high correlation by predicting the global average expression rather than perturbation-specific responses, which yields false positives and poor interpretability. Methods that add biological knowledge graphs typically treat them as dense, static priors shared across perturbations, propagating noise. We propose AdaPert, which counters mean collapse by extracting a sparse, perturbation-specific subgraph via differentiable node selection, then suppressing spurious variation in non-responsive genes while emphasizing differentially expressed ones. Across multiple benchmarks, AdaPert outperforms existing baselines, with the largest gains on DEG-aware metrics.
Lay Summary
Cells contain thousands of genes that work together. To learn what a gene does, scientists switch it off and observe how the cell reacts—a technique central to drug discovery. Because testing every gene in the lab is impossible, researchers use AI to predict the outcomes. Current models often take a shortcut: they predict an average response that looks similar regardless of which gene was changed, missing the few genes that actually matter. AdaPert avoids this by using biological knowledge to focus on the small group of genes most likely to be affected for each specific change. It identifies the truly responding genes more accurately than existing methods, making AI predictions more useful for biological research and drug discovery.