Structural Locality Differentiates Residue-Tokenised Bidirectional and Causal Protein Language Model Families
Abstract
Protein language models (PLMs) trained under different bidirectional and causal regimes achieve strong downstream performance, yet how these architectural and objective differences jointly shape learned representations remains poorly understood. We apply TopK sparse autoencoders (SAEs) to three residue-tokenised PLMs (ESM-2, RITA, ProtT5; the ProtT5 encoder and decoder are probed separately) at nine matched relative depths. We find that ESM-2 (bidirectional) learns features with stronger structural locality than RITA (residue-tokenised causal) at every matched depth, robust across metric settings, held-out proteins, and SAE initialisations. On a BPE-tokenised PLM (ProtGPT2), we identify a tokenisation artefact in residue-pair locality metrics—a large naïve sequential-locality effect that reverses under inter-token control—motivating our restriction to residue-tokenised models in the main analyses. Within ProtT5, the nine-depth grid localises a structural-locality crossover between encoder and decoder to ≈42% relative depth, consistent with cross-attention propagating encoder-derived context that fades under the autoregressive objective. Within-model depth trajectories for ESM-2, RITA, ProtT5-enc, and ProtT5-dec are qualitatively distinct. These findings support a single robust bidirectional-vs-causal dissociation (on structural, not sequential, locality) and identify a methodological interaction between standard residue-projection and residue-pair locality metrics on BPE-tokenised PLMs.