Decoding Loss-of-Function Variants with Sparse Concept Features of ESM-2
Abstract
Protein language models (PLMs) excel at variant-effect prediction, yet their dense residue embeddings lack interpretable axes, making the basis for predictions unclear. To address this, we trained sparse autoencoders (SAEs) on wild-type embeddings and aligned the resulting features with functional residue annotations, yielding interpretable concept features—SAE dimensions aligned with biological annotations (e.g., kinase domain, ATP binding site). Without any deep mutational scanning (DMS) supervision, we found that the more residue positions where a missense variant silences a concept feature below its wild-type activation, the greater the loss of function (LOF) observed in DMS assays. Furthermore, across five kinase activity assays, suppression of these concept features localized to subdomain IX—a region whose disruption is known to substantially impair kinase activity. These findings show SAEs decompose dense language model representations into interpretable signals that align with known mechanisms underlying variant effects.