Toxin Feature Hierarchy in ESM-2: Mechanistic Interpretability reveals Why Frozen Probes Resist ProteinMPNN Redesign
Manan Wadhwa ⋅ Shivam Dubey
Abstract
Sequence-identity screening (BLAST with $e$-value $\leq 10^{-3}$ and identity $\geq 40\%$) fails entirely against ProteinMPNN redesigns: 0\% detection across 643 redesigns below the 40\% identity threshold. A linear probe on frozen ESM-2 650M embeddings detects 93.9\% of a 534-redesign test subset with no exposure to redesigned sequences during training.This is a $+$93.9 percentage-point gap explained mechanistically. Using interPLM Sparse Autoencoders (SAEs), we identify a 50-feature set (${\sim}38{\times}$ compression from active features) whose mean transfer ratio of \textbf{1.28} shows that ProteinMPNN \emph{amplifies} toxin structural features, because it preserves the backbone topology the feature set encodes. Direct Probe Attribution (DPA) pinpoints layer 32 as the detection bottleneck ($r{=}0.992$ redesign--toxin feature correlation; 65\% feature overlap). A four-tier attack taxonomy places the security boundary precisely at gradient access: 6.1\% evasion (blackbox ProteinMPNN) vs.\ 100\% (white-box gradient). SAE probes recover 38\% of ``Double-Evaders'' that fool both BLAST and the linear probe, demonstrating direction-sensitive detection beyond Euclidean boundaries. Zero-shot scanning of UniRef50 reveals generalization beyond training distribution: 248 candidates show cross-kingdom transfer (23 fungi despite zero training), structure-agnostic detection (pLDDT-independent), and 4.75${\times}$ signal-peptide enrichment, suggesting the probe learned how to generalize.54\% uncharacterized, 31 from WHO Priority-1 pathogen
Chat is not available.
Successful Page Load