Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Abstract
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors. Current methods often rely on costly LLM-based techniques (e.g. annotating dataset differences) or dense embedding models (e.g. for clustering), which lack control over the properties of interest. We propose using sparse autoencoders (SAEs) to create SAE embeddings: representations whose dimensions map to interpretable concepts. Through four data analysis tasks, we show that SAE embeddings are more cost-effective and reliable than LLMs and offer the controllability that dense embeddings lack. Using the large hypothesis space of SAEs, we can uncover insights such as (1) semantic differences between datasets and (2) unexpected concept correlations in documents. For instance, by comparing model responses, we find that Grok-4 clarifies ambiguities more often than nine other frontier models. Relative to LLMs, SAE embeddings uncover bigger differences at 2-8x lower cost and identify biases more reliably. Additionally, SAE embeddings are controllable: by filtering concepts, we can (3) cluster documents along axes of interest and (4) outperform dense embeddings on property-based retrieval. Using SAE embeddings, we study model behavior with two case studies: investigating how OpenAI model behavior has changed over time and finding "trigger" phrases learned by Tulu-3 (Lambert et al., 2024) from its training data. These results position SAEs as a versatile tool for unstructured data analysis and highlight the neglected importance of interpreting models through their data.
Lay Summary
Sparse autoencoders (SAEs) decompose large language model (LLM) activations into interpretable features e.g. "dog", "cat", "positive tone". Thus, when text is run through an LLM, we can use the SAE to capture information found in the LLM's activations, essentially labelling the text with thousands of features at once. This creates an "interpretable embedding" of text which we can use in many different ways---for example, finding features that differ between datasets, finding feature correlations in data, clustering data using controllable features, or retrieval data with certain features. We show how these can be useful especially for analyzing language model outputs and training data. For example, using features in large chat transcript datasets, we find that Grok 4 surprisingly expresses more caution around ambiguous queries than other frontier models.