scCBGM: Single-Cell Editing via Concept Bottlenecks
Abstract
Understanding cellular phenotypes and how they respond to perturbations is critical for disease biology and therapeutic design. Single-cell RNA sequencing enables characterization at cellular resolution, yet the combinatorial space of conditions makes exhaustive experimental mapping infeasible. We introduce single-cell Concept Bottleneck Generative Models (scCBGM), a framework for interpretable and precise counterfactual editing of individual cells. scCBGM adapts concept bottleneck architectures for single-cell data through decoder skip connections and a cross-covariance penalty that promotes disentanglement without dimensional constraints. We extend the framework to flow matching models, enabling concept-guided editing in both encoding-decoding and generation regimes. To enable rigorous evaluation, we develop a synthetic benchmark with ground-truth counterfactuals. Across multiple real datasets, scCBGM demonstrates superior performance in combinatorial generalization and counterfactual prediction, supported by cell-level validation on synthetic data and population-level benchmarks on real datasets.
Lay Summary
Problem: Understanding how individual cells respond to factors like drugs or diseases is vital for developing new drugs and therapeutics. However, because there are millions of different experimental conditions to test, running every single experiment is impossible. We therefore need computational methods that allow us to work with limited experimental data and generalize to unseen combinations of relevant conditions (e.g., cell type, drug, and dose). Solution: To tackle this, we developed a new generative framework called single-cell Concept Bottleneck Generative Models (scCBGM). This framework acts like an "editing tool" for cells, letting us modify different variables to ask questions like “what if A instead of B” and observe the resulting effect on the cell. This makes it possible to, for example, apply drugs to new cell types or activate a pathway in a specific cell and study the outcome computationally. Impact: This work could allow scientists to significantly scale the number of hypotheses they test, giving them an estimate of what changes given interventions would produce. These computationally experiments are never meant to replace real wet lab experiments, but to augment them and narrow down the list of experiments that are most likely to succeed.