The Dark Subspace of Fine-Tuning Memorisation
Abstract
Sparse autoencoders (SAEs) learn a dictionary of sparse features over neural-network activations and are widely used to interpret and edit language models. Privacy and safety methods now ablate or steer these features to suppress unwanted behaviour, acting only on what the dictionary represents. We ask whether the dictionary preserves the evidence that a document was used to fine-tune the model, the question studied by membership-inference attacks. We decompose each activation into the SAE reconstruction and the reconstruction residual, then train membership detectors on each. In controlled Pythia replications, SAE reconstruction weakens detection, yet detectors recover much of the lost signal from the residual. The same residual-above-reconstruction ordering holds across ten model-SAE settings spanning six architecture families. Surface confounds (norm, length, bag-of-words) and a label-shuffle permutation test do not account for the pattern. Privacy and safety methods that operate only on SAE features may therefore leave training-data evidence detectable in the reconstruction residual.