bish-bash-fold: what are protein structure prediction models learning?
Abstract
Protein structure prediction models rely on coevolutionary signals, intrinsically limiting their efficacy for de novo protein design. Here, we investigate this failure regime by training sparse autoencoders on the diffusion modules of AlphaFold3 and Boltz-2 to recover interpretable features along the denoising trajectory. By ablating MSA depth and sequence input, we categorise features into MSA-activated, MSA-silenced, sequence-driven, and model-prior classes. We find that AlphaFold3 is less sensitive to input degradation while Boltz-2 is more conditioning-dependent. Using the \citet{garcia2025evaluating} dataset, we then investigated whether these features could provide mechanistic insight to designability failure modes. Linear probes trained on feature activations predict experimental outcome (measured by expression, solubility, monomericity, and CD) with higher accuracy and precision than pLDDT, and recover known design heuristics. As the first interpretability study of structure prediction denoising trajectories, this work establishes that internal representations capture critical designability signals beyond standard model outputs.