Scaling Laws for SAE Training Data
Nikita Koriagin ⋅ Nikita Balagansky ⋅ Daniil Gavrilov
Abstract
Sparse autoencoder (SAE) training is often bottlenecked by activation storage. We show that the stored activation buffer can be far smaller than the total SAE training budget: quality depends mainly on the number of \emph{unique} activations, not on whether every training token is fresh. After a short diversity-limited regime, additional fresh activations provide little benefit; replaying the same buffer preserves both reconstruction and interpretability performance. We capture this diversity--repetition tradeoff with a data-constrained scaling law and validate it across dictionary sizes, token budgets, Llama-3.1 and Qwen3 models, early/middle/late layers, BatchTopK/TopK/JumpReLU SAEs, and downstream metrics including EV, reconstruction MSE, feature absorption, automated interpretability, and sparse probing. The resulting buffer-sizing rule is simple: choose the smallest activation buffer that reaches the quality plateau, then use replay to spend the remaining training budget. In our experiments, this reduces activation storage by $8\text{--}64\times$ with negligible quality loss.
Chat is not available.
Successful Page Load