ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation
Abstract
Lay Summary
Proteins that bind tightly and specifically to a chosen target molecule — "protein binders" — are central to modern medicine, diagnostics, and biotechnology. Recently, AI methods have made it possible to design such binders from scratch computationally, generating thousands of candidate designs in a short time. But before any candidate is tested in the lab (which is slow and expensive), researchers must first judge computationally which designs are likely to work. The problem is that there is no agreed way to do this: different studies use different AI "verifier" models, different cut-offs, and different definitions of success, so their reported results cannot be fairly compared. We introduce ProtDBench, a standardized framework for evaluating protein binder design. Using a large set of designs with real laboratory results, we show that the choice of verifier strongly shapes which designs appear successful, and that no single verifier is reliable on its own. We then compare leading open-source design methods on the same footing, and show that judging a method by accuracy alone is misleading, generation speed and structural diversity matter just as much, and these properties often do not go together. ProtDBench gives the community a fair, reproducible way to compare protein-design methods, helping focus laboratory effort on the most promising candidates.