Feature Recovery Requires Structured Event Regimes in Sparse Reconstruction
Abstract
Sparse autoencoders are often used in mechanistic interpretability as if sparse reconstruction should recover the features represented by a model. Recent work shows that this recovery is fragile, but it remains unclear which failures come from the SAE architecture, the encoder, optimization, or finite data. We show that several failures can already be incentivized by the population-level sparse-reconstruction objective. We study by how much residual mass projects above the sparsity threshold in a positive linear latent-ray model; merging and splitting arise as static properties of this objective, while absorption and seed-dependent alternatives arise sequentially as earlier selections change the residual field. We also separate two notions often conflated in interpretability practice: recovering a ground-truth direction and recovering the activation pattern of that feature. Neither implies the other in general; they coincide under specific structural conditions, such as single-feature event dominance and regular simplex structure in learned symmetric geometries. Sparse reconstruction therefore recovers ground-truth features only in structured event regimes; outside them, the objective can favor non-canonical but reconstruction-useful directions, and a one-ReLU encoder introduces a further representability gap governed by whether the oracle gate is affine-ReLU approximable. Overall, our results refine the existing analysis of SAE behavior and provide a unified perspective on ground-truth feature recovery studies.