MutAtlas: A PDB-Wide Energy-Guided Atlas of Protein Mutation Effects
Abstract
Protein mutation effect prediction is fundamental to protein engineering and disease variant interpretation, yet experimentally measured mutation data remain accurate but extremely sparse. To provide scalable supplementary mutation signals, we construct a PDB-wide mutation augmentation dataset that exhaustively enumerates single-site substitutions on experimentally resolved protein structures and aligns mutation signals from physics-based energy models, protein language models, and inverse folding models. Large-scale analysis under a unified mutation preference representation reveals substantial differences in the consistency, concentration, and substitution patterns of mutation distributions across models, indicating that disagreement is pervasive and reflects conflicting inductive biases rather than random noise. Motivated by these observations, we propose an unsupervised multi-source mutation preference distillation framework that learns from relative mutation preferences while explicitly modeling cross-source disagreement. Without using any experimental mutation labels during training, our approach achieves the best overall performance among the evaluated zero-shot baselines and naive multi-source fusion strategies on ProteinGym. We release the dataset and evaluation pipeline to support reproducible studies of protein mutation effects.
Lay Summary
Understanding how changes in a protein sequence affect its behavior is important for designing useful proteins and interpreting disease-related genetic variants. However, reliable experimental measurements are available for only a small fraction of possible protein mutations. In this work, we build a large resource that uses existing protein structures to generate computational signals for many possible single-position mutations. These signals come from different types of models, including physics-based energy calculations and modern protein AI models. We find that these models often disagree about which mutations are favorable, and that this disagreement is not just random error. Instead, different models capture different aspects of proteins. Based on this finding, we develop a learning method that combines information from multiple sources while also accounting for where they disagree. Without using experimental mutation measurements during training, our method improves overall performance on a standard protein mutation benchmark. We will release the dataset, code, and evaluation pipeline to support future research in protein engineering and mutation effect prediction.