Mechanistic Interpretability of Chemical Language Models: Molecular Geometry and Representations in MoLMFormer and SMI-TED
Abstract
Chemical language models trained on SMILES strings, text-based representations of molecules, achieve strong performance on chemical prediction tasks, yet it remains unclear how they internally represent molecular structure and chemical properties. In this work, we analyze two widely used models, MoLMFormer and SMI-TED, from a mechanistic interpretability perspective. Prior analyses of chemical foundation models have largely focused on correlational methods, including attention correlations, embedding structure, and input-output associations. We take a step toward a more mechanistic understanding by combining layerwise attention analysis, linear probing, and causal intervention. We find that attention exhibits only weak and selective alignment with underlying molecular geometry. Probing further reveals that SMI-TED encodes contextual chemical properties as cleaner linear directions than MoLMFormer. Moving beyond correlational analysis, our causal interventions show that these directions are functionally used during computation, and earlier in SMI-TED than in MoLMFormer. Despite strong downstream performance, our results suggest that successful SMILES pretraining does not necessarily produce spatially grounded or uniformly usable chemical representations, highlighting how diverse internal representations can support similar predictive success.