Faithful Is Not Interpretable: Sparse Features, Circuits, and Robustness in Frozen Audio Encoders
Arnesh Batra ⋅ Aniket Khandelwal ⋅ Arush Gumber ⋅ Krish Thukral
Abstract
Audio ML systems increasingly reuse frozen encoders for event classification, music, speech affect, ASR-related transfer, and audio-text retrieval. A linear probe can show that a representation works, but not whether its evidence is robust under acoustic shift, concentrated in interpretable features, faithfully routed across layers, or safe to edit. We audit seven frozen audio encoders with sparse autoencoders, sparse transcoders, controlled corruptions, steering, and feature ablation. Two separations dominate. First, faithful inter-layer prediction is not interpretability: transcoders explain up to $R^2=.998$ of the next-layer representation while Whisper-S has zero class-monosemantic routing, and the correlation between explained variance and monosemanticity across stored transitions is only .123. Second, global stability is not feature stability: Whisper-S keeps directionally stable embeddings under corruption, yet its classifier-relevant sparse features stop firing, producing the largest score drop. By contrast, CLAP-HTSAT concentrates label evidence into about eight SAE features on average, while speech encoders distribute evidence over broader, more polysemantic sets. These diagnostics help audio practitioners distinguish compact acoustic detectors from robust distributed codes and faithful dense routes before deploying or editing frozen audio models.
Chat is not available.
Successful Page Load