LipoPU: Pocket-level Prediction of Lipid-Protein Interactions via Positive-Unlabeled Learning
Abstract
Computational identification of lipid-binding proteins is critical for both fundamental research and therapeutic development. Existing models are typically trained in a fully supervised manner, treating unlabeled samples as negatives. However, missing evidence does not imply non-binding, leading to systematic false negatives. Pocket-level lipid-binding prediction also remains underexplored compared to residue- or sequence-level approaches. To bridge these gaps, we present LipoPU, a pocket-centric predictor that formulates lipid-binding learning under a ranking-based positive-unlabeled objective, and supports both binary lipid-binding detection and multi-label lipid category prediction. LipoPU learns an attention-based pocket representation that is robust to ambiguous pocket definitions while providing residue-level interpretability. Experiments show consistent gains over supervised baselines and prior pocket-level work, and a structural case study recovers a literature-supported allosteric lipid-binding pocket while highlighting biologically informative residues.
Lay Summary
Lipids play central roles in many cellular processes, but experimental observations of lipid–protein interactions remain limited and sporadic. In this work, we propose LipoPU, a machine learning method that learns from confirmed lipid-binding examples while recognizing that unreported cases may still include real lipid–protein interactions. LipoPU focuses on protein pockets, local regions where molecules such as lipids bind to proteins. Using a ranking-based training strategy, LipoPU gives higher scores to candidate pockets with stronger lipid-binding potential. It goes beyond asking whether a pocket may bind lipids and predicts broad lipid categories, such as sterols or glycerophospholipids. Compared with conventional machine learning methods that treat unreported cases as non-binding, and with prior pocket-level work, LipoPU performs better on both proteins closely related to the training data and proteins with remote homology. By assigning attention weights to pocket residues, it highlights residues that provide stronger clues for predictions, making the results easier to interpret biologically. By helping researchers prioritize promising pockets, LipoPU turns incomplete lipid-binding records into a practical starting point for large-scale screening and target discovery.