KERMT: Improving Graph Foundation Models for Drug Discovery through Multi-task Fine-Tuning, Data Scaling, and Accelerated Training
Abstract
Chemical foundation models can improve prediction of critical properties in small molecule drug discovery by leveraging representations learned from self-supervised pretraining. We implement multi-task fine-tuning in two pretrained graph transformer models, KPGT and KERMT (Kinetic grovER Multi-Task), our enhanced version of the GROVER model. We benchmark the models on diverse on-target biology assay and off-target ADMET (absorption, distribution, metabolism, excretion, and toxicity) prediction tasks, and show that KERMT has improved predictive performance for ADMET endpoints with 2,000 or more data points for training. Contrary to popular hypotheses that pretrained models show better performance in data-scarce regimes, KERMT improves most as dataset size increases. KERMT generalizes better than Chemprop across all compound similarity bins, while additional pretraining on chemical structures more similar to the downstream structures does not improve performance. To support future work, we release two multi-task ADMET splits and an accelerated implementation of KERMT for scalable pretraining, fine-tuning, and inference.