Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability
Rodrigo Mendoza Smith
Abstract
Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting an overcomplete dictionary $\mathbf{W}\in\mathbb{R}^{m\times n}$ with $m
Chat is not available.
Successful Page Load