Hepa-RAFT: Retrieval-Augmented Virtual Hepatocyte Responses for Hepatotoxicity Prediction
Abstract
Generative and agentic systems for therapeutic design need inexpensive biological feedback before prioritizing compounds for synthesis or wet-lab testing. Drug-induced liver injury (DILI) is a high-impact case: structure-only predictors capture chemical liabilities but miss cellular response, whereas measured transcriptomic assays provide mechanism but are too costly for early-stage screening. We present Hepa-RAFT, a Hepatocyte Retrieval-Augmented Framework for Toxicogenomics that provides a structure-conditioned virtual hepatocyte-response readout: it maps a query SMILES to predicted drug-induced hepatocyte expression and uses that model-derived state as an intermediate retrieval key. The framework has two task-adaptive heads sharing one FiLM-conditioned expression predictor: an expression head corrects predicted profiles by retrieving training drugs with similar predicted perturbation, and a DILI head performs similarity-weighted voting over a small labeled reference set using chemical, predicted-transcriptomic, pharmacokinetic, and drug--target channels. This yields a biological design principle: retrieval should be matched to the phenotype being queried, with perturbation correction favoring a focused predicted-response geometry and DILI liability favoring broader evidence aggregation. Across held-out scaffold splits, the end-to-end retrieval-augmented Hepa-RAFT pipeline reaches Pearson correlation 0.631 on the top 20 perturbation-responsive genes; on external DILImap and DILIst benchmarks, Hepa-RAFT reaches AUROC 0.713 and 0.637 without retraining. Because query-drug expression and labels are never accessed, Hepa-RAFT is best interpreted as a lightweight, biologically grounded virtual-response module for hepatocyte-intrinsic DILI liability, with training-overlap-controlled external comparisons and overlap audits reported for available competitor training sets.