PDFBench: A Benchmark for De Novo Protein Design from Function
Abstract
Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering. However, the field lacks a unified and comprehensive evaluation framework. Current models are assessed using inconsistent and limited subsets of metrics, which prevents fair comparison and a clear understanding of the relationships between different evaluation criteria. To address this gap, we introduce PDFBench, the first comprehensive benchmark for function-guided de novo protein design. Our benchmark systematically evaluates eight state-of-the-art models on 16 metrics across two key settings: description-guided design, for which we repurpose the Mol-Instructions dataset, originally lacking quantitative benchmarking, and keyword-guided design, for which we introduce a new test set, SwissTest, created with a strict datetime cutoff to ensure data integrity. By benchmarking across a wide array of metrics and analyzing their correlations, PDFBench enables more reliable model comparisons and provides key insights to guide future research.
Lay Summary
Proteins are tiny biological machines that carry out many essential jobs in living organisms. Designing new proteins with desired abilities could help create better medicines, industrial enzymes, and research tools. However, it is hard to compare today’s AI systems for protein design because different studies use different tests, making it unclear which systems work best and why. This paper introduces PDFBench, a benchmark for evaluating AI models that design proteins from a requested function. The request can be written as a plain-language description or provided as standard biological function labels. PDFBench compares eight recent protein design models using sixteen tests that check whether the designed proteins look realistic, are likely to fold into usable shapes, match the requested function, differ from known proteins, and offer diverse design choices. The benchmark also includes a newly built test set designed to reduce the chance that models have already seen the answers. Our results show that some models produce more realistic and better-folding proteins, while others explore more diverse designs. We also find that some commonly used tests can be misleading unless interpreted carefully. Overall, PDFBench provides a fairer and more complete way to measure progress in AI-based protein design.