Phase-Calibrated Steering of Protein Diffusion Language Models
Abstract
Test-time steering of protein generative models has emerged as a leading paradigm for guided sequence design, with applications ranging from clinical variant interpretation to therapeutic protein engineering. Contemporary approaches, notably Twisted Sequential Monte Carlo and ensemble-guided sampling over diffusion language models such as DPLM-650M, have achieved solid generative performance by tilting the model with a fixed reward learned end-to-end from multi-mutant fitness data. However, these methods typically treat reward calibration as a learned object and ignore the iid-sum structure of additive fitness landscapes. On the multi-mutant benchmark a simple sum-of-singles additive predictor outperforms every published learned model on 115 of 116 proteins. To build a better calibrated steering procedure, we recognize edit distance as a survival time and the per-protein viability function as the survival curve of an iid sum of single-mutant effects drawn from the protein's DMS spectrum. Based on this, we introduce PhaseSMC, a phase-calibrated Twisted Sequential Monte Carlo framework for protein editing that uses the closed-form survival prior as a calibrated reward, and requires no multi-mutant labels for new proteins. Our novel framework is consistent and can be adapted to any masked-LM or diffusion backbone. Together, these advances lay the groundwork for label-efficient, calibrated generative protein design at proteome scale, with immediate applications in clinical variant interpretation and therapeutic protein engineering, and broader opportunities across generative-modelling domains in which the underlying data are well-approximated by iid sums.