Virtual Cell Models Inflate Perturbation Effect Sizes and Undermine Causal Gene Regulatory Network Recovery
Aayan Alwani ⋅ Ethan Y Wang
Abstract
Virtual cell models are deep neural networks that predict transcriptional responses to genetic perturbations and are proposed as in silico substitutes for laboratory CRISPR screens. Ahlmann-Eltze and Huber (2025) recently showed that current models fail to outperform simple linear baselines at per-gene prediction; we ask whether the same holds downstream when their predictions are used to recover causal gene regulatory networks. We introduce PerturbCausal, a benchmark that pushes predictions from four virtual cell models (GEARS, CPA, Geneformer, and the 2025 Arc Institute STATE foundation model) through four causal discovery algorithms (GIES, inspre, PC, NOTEARS) on the Replogle K562 and RPE1 Perturb-seq datasets and the Norman 2019 K562 dataset. Across five seeds, the three pre-2025 models do not significantly outperform a sparse linear regression trained on control cells alone. STATE is the first model in the roster to materially exceed the linear baseline ($F_1 = 0.462$), but its absolute causal F1 remains below $0.5$, leaving more than half of real interventional edges unrecovered. We reconcile two competing narratives in the recent literature, mode collapse and effect-size inflation, as two views of the same failure: deep models predict similar dense response patterns across perturbations. A model-agnostic quantile-matching calibration corrects this magnitude failure exactly but recovers only about 40% of the causal gap, locating the remainder in joint-distribution structure that distributional calibration cannot reach. The benchmark, calibration toolkit, and reproducibility code are released as an open-source Python package.
Chat is not available.
Successful Page Load