ProMiSE: Protein Multi-State Evaluation Benchmark in Biological Contexts
Abstract
Proteins are inherently dynamic, with biological functions often emerging from transitions between multiple conformational states. While recent breakthroughs have largely addressed the static structure prediction problem, no systematic benchmark exists to demonstrate how well current models capture functionally relevant dynamics. We introduce ProMiSE, the first benchmark that provides both a dataset and an evaluation scheme, based on native biological assemblies and integrating major conformational change mechanisms—intrinsic, ligand-induced, and protein-induced—within a single curated dataset. We conducted a comprehensive evaluation of state-of-the-art structure prediction models, including AlphaFold3 and recent generative approaches. Our findings reveal that current models exhibit a limited ability to sample intrinsic multi-states and are often insensitive to biological context in induced scenarios. Internal representation analysis suggests that training-data exposure can shift predictions toward dominant conformational states over alternative biologically relevant states, primarily at the structure module. In contrast, results from BioEmu indicate that reducing decoding-stage bias can substantially improve multi-state sampling without major changes to upstream pair representations.
Lay Summary
Proteins are tiny molecular machines, and many of them work by shifting between several different shapes—for example, one shape to grab a target and another to release it. Recent AI tools have become remarkably good at predicting a protein's structure, but they usually guess just one shape, leaving open how well they capture this shape-shifting that drives biology. We built ProMiSE, a carefully curated test set and scoring system that checks whether leading AI models—including AlphaFold3 and newer methods—can recover the multiple shapes a protein is known to adopt, across the main biological reasons proteins change shape. We found that today's models tend to "collapse" onto a single favored shape and miss the alternatives, often ignoring the biological context that should trigger a change. By tracing where this failure happens inside the models, we showed the bottleneck lies in the final shape-building step, where the model leans on shapes it saw most often during training rather than exploring others. Encouragingly, one model that was retrained only at this final step sampled diverse shapes far better. ProMiSE gives the field a shared yardstick and a concrete direction for building AI that understands proteins as the dynamic machines they really are.