PerturbReason: A Knowledge-Grounded Benchmark and Framework for Cell-State–Conditioned Mechanistic Reasoning of Perturbation Effects
Abstract
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under distribution shifts. We introduce PerturbReason, a knowledge‑grounded benchmark for cell‑state--conditioned reasoning about perturbation effects. It tests whether models can generate mechanistically faithful explanations and assesses their robustness against complex shifts, such as new cells and unseen perturbations. PerturbReason combines multiple single‑cell genetic and chemical perturbation data with knowledge graphs, and dynamically conditions pathways on cell‑specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between outcome prediction and mechanistic reasoning, exhibiting failure modes invisible to standard benchmarks. As a reference probe, we present PerturbRM, a large language model that aligns outcome predictions with context-specific regulatory reasoning. PerturbReason thus provides a rigorous diagnostic benchmark for studying faithful, generalizable reasoning in data-rich scientific systems.