TSMGen: Target-Specific Molecule Generation via Higher-Order Structural Dependencies and Context-Aware Bidirectional Fusion
Abstract
Lay Summary
Modern drug discovery is expensive and time-consuming, especially when scientists need to design molecules that can precisely bind to disease-related proteins. Artificial intelligence has recently shown promise in helping generate new drug candidates, but many existing methods still struggle to fully understand the complex structure of protein binding sites. In particular, they often focus only on simple pairwise relationships between amino acids and overlook more complicated structural interactions that are important for molecular recognition. In this work, we introduce TSMGen, an AI framework for target-specific molecule generation. Our method studies protein binding pockets at multiple levels, capturing both detailed local atomic information and broader structural patterns across groups of amino acids. It also allows the protein and the generated molecule to continuously exchange information during the design process, helping the model produce molecules that better match the target protein environment. We evaluated our approach on public drug discovery datasets and compared it with several state-of-the-art molecular generation methods. The results show that TSMGen generates molecules with stronger predicted binding ability, better drug-like properties, and higher chemical quality. In a case study involving a protein related to Alzheimer’s disease, our method produced candidate molecules with stronger predicted binding affinity than known reference compounds. These results suggest that TSMGen could help accelerate the early stages of drug discovery by improving the design of candidate molecules for specific disease targets.