Large-scale sequence modeling of antibody-antigen binding specificity
Abstract
Antibody-antigen binding specificity underlies immune protection, vaccine efficacy, and therapeutic antibody development, yet remains difficult to predict. Existing approaches either rely on structure prediction methods ill-suited to antibody-antigen complexes, or sequence-based protein-protein interaction models that require inter-protein coevolution signals absent in antibody-antigen systems. To address this, we develop Agate, a discriminative protein language model trained on over one million human antibody-antigen pairs curated from public datasets, the largest dataset to date. Agate jointly encodes antibody-antigen pairs and is trained using both experimentally validated binders and biologically grounded negative, non-binder examples. To rigorously assess generalization, we implement stringent sequence-based train-test partitioning, including evaluation on completely unseen antigens. Agate achieves state-of-the-art performance for antibody-antigen binding prediction while substantially improving specificity, scalability, and prospective generalization relative to existing methods. Together, these results establish a scalable framework for modeling of antibody recognition with applications in pandemic preparedness, vaccine design, and precision immunotherapy.