Flexible Kernels for Protein Property Prediction
Abstract
Despite its importance to applications in protein design, predicting protein properties like binding affinity and thermostability from sparse experimental data remains a significant challenge. Accordingly, we introduce a class of sequence kernels that exploit evolutionary substitution matrices as well as local linearity and demonstrate that the resulting Gaussian processes provide data-efficient models of protein property landscapes, frequently outperforming alternatives that rely on foundation model embeddings. Furthermore--by learning what are in effect structure-aware substitution matrices--we show that our kernels can readily incorporate structural information from foundation models. We demonstrate that these structure-conditioned kernels are well suited to multi-task learning across multiple protein property landscapes and can decisively outperform local supervised learning methods.
Lay Summary
Protein design offers great potential to create new and improved medicines, industrial enzymes, and tools for biological research. Realizing this potential remains challenging, however, because laboratory experiments can test only a limited number of proteins. To help address this, we developed a machine learning method that predicts how changes to a protein sequence will affect its properties, such as binding strength or stability. These predictions let researchers narrow down which proteins to test in the lab, speeding up the protein design process. The method uses knowledge from evolution to make accurate predictions from limited data and can incorporate 3D protein structure information when it is available. Across many benchmarks, our approach often matches or outperforms more complex methods that rely on large AI models trained on massive protein datasets, while being simpler and more data-efficient.