FluorCode: Predicting Fluorescent Protein Photophysical Properties with LoRA-Fine-Tuned Protein Language Models
Rico C. K. Sou ⋅ Alicja Ziajowska
Abstract
Predicting fluorescent protein (FP) photophysical properties from sequence remains challenging, especially for novel fluorescent proteins because of the limited diversity of existing datasets. Using FluorCode, a fine-tuned protein language model for photophysical properties, we compare alignment-based one-hot features against LoRA-fine-tuned ESM2-650M with chromophore-aware attention pooling. Under sequence-identity-clustered cross-validation at 50\%, the alignment-based baseline accuracy degrades sharply, with excitation MAE increasing from 12.7 to 35.6\,nm and brightness prediction even collapsing to noise ($R = 0.13$), whereas the LoRA-ESM2 MLP model remains robust (excitation MAE: 8.9 $\rightarrow$ 11.3\,nm; brightness $R = 0.96$). Meanwhile, adding Pocket-3D chromophore-anchored descriptors from 913 chromophore-grafted structures yields no consistent measurable gain beyond the sequence model, indicating that explicit structural information in this hand-crafted form does not improve over LoRA-ESM2 under our current setup. Here, we demonstrate that the standard random cross-validation protocol used in prior work, previously shown to inflate performance in protein machine learning but not addressed in fluorescent protein prediction, overestimates performance by placing near-identical FP variants in both training and test folds.
Chat is not available.
Successful Page Load