Auditing a Multi-Modal Chromatin Foundation Model with Sparse Autoencoders
Abstract
Multi-modal foundation models for biology are increasingly deployed in life sciences pipelines, yet what they internally represent remains poorly understood. We audit EpiBERT, a transformer that jointly processes DNA sequence and ATAC-seq chromatin accessibility, asking whether it encodes a clinically consequential contrast: in vitro (cell line) vs. in vivo (primary tissue) chromatin context. We train layer-wise Sparse Autoencoders(SAEs) with BatchTopK activations across six matched ATAC-seq conditions, introduce the Context Divergence Score (CDS) to identify context-specific features, and validate them via causal ablation, linear context-steering, and three-level biological annotation. Context-specific features grow 3.8-fold from early to late layer; causal ablation yields a large effect (Cohen’s d=1.79); context- steering closes 11.2% of the prediction gap at 4.5× above random; and discovered features are enriched for lineage-defining transcription factors (HNF4A/FOXA2 in liver, SPI1/RUNX1 in blood, EBF1/PAX5 in lymph). These results provide an interpretability and robustness audit of a multi-modal biological foundation model and a concrete intervention path for tissue-aware deployment.