Interpreting Genomic Language Models using Sparse Autoencoders
Abstract
Genomic language models (gLMs) achieve strong performance across genomic prediction tasks, but their internal biological representations remain poorly understood. Sparse autoencoders (SAEs) have emerged as an interpretability tool in vision and natural language models, yet their applicability to gLMs remains unexplored. We present a systematic study of SAE-based interpretability for gLMs, introducing a diverse benchmark of human genomic annotations and a suite of genome-tailored interpretability metrics. Using Evo2 as a primary case study, we show that SAE features, particularly those from intermediate layers, are more interpretable than raw model embeddings across 42/55 (76%) of our genomic concept evaluations, with 26 of them having an F1 score greater than 0.7. We further find that interpretability depends on SAE training data properties such as evolutionary proximity and context length. Finally, to organize semantically related genomic concepts learned by an SAE, we develop a graph-based representation method that outperforms the baseline approach of using SAE model weights. We demonstrate how our framework can extend SAEs as a powerful approach for not only better understanding gLMs but also for adopting them in disease-driven genomic explorations.
Lay Summary
Genomic language models (gLMs) - artificial intelligence (AI) models trained on DNA sequences - are becoming increasingly powerful at predicting biological function. What can we learn from these models and how can we make them more transparent and easier to interpret? In this study, we investigated whether sparse autoencoders (SAEs), an AI interpretability method, could help uncover the biological information hidden inside gLMs. Using a gLM called Evo2 as a case study, we developed new ways to test whether the patterns identified by SAEs correspond to meaningful biological concepts. We found that SAEs revealed clearer and more interpretable biological signals than the original model representations, including signals linked to numerous important functional properties of DNA. We also created a “feature atlas” that organizes related biological concepts into an intuitive map of what the model has learned about DNA. Using this atlas, we can visualize how disease-associated genetic variants can alter DNA grammar and function through the lens of the AI model. By making genomic AI models easier to interpret, this work can help us better leverage them to understand genetic variation and the biological basis of human disease.