Evaluating Sparse Autoencoders in a Structured Continuous Model: Concept Recovery in Implied-Volatility Forecasting
Abstract
Standard SAE concept-recovery protocols rely on correlation alignment, top-K stability, and patching point estimates without explicit nulls. We propose matched-random baselines (leave-one-out random latents for single-latent patching, and matched-random latent pairs for interaction patching, both with bootstrap confidence intervals) and apply them in a structured continuous testbed where ground-truth descriptors are available: a compact MLP trained to forecast SPY implied volatility from tick-level NBBO data over 45 trading sessions. The standard pipeline reports five cross-seed-stable concept families (moneyness geometry, lagged skew, spot level, rates, lagged curvature), and ridge probes from the SAE codes recover four non-geometric descriptors (lagged skew, lagged curvature, spot, lagged ATM level) at positive held-out R² (0.27–0.49). Under matched-random patching, however, only one family survives: moneyness geometry produces single-latent effects at 4.49× its leave-one-out baseline (p<0.001), and the moneyness × lagged skew interaction clears the matched-random null (p=0.030, N=300 random pairs). Every other non-geometric family's mean |Δ| sits at or below its baseline. The probe-patching gap is the paper's central methodological observation: concepts that the SAE encodes (probes) and concepts that the model uses (patching) are distinct populations. We argue matched-random baselines should be standard for SAE-interp validation, and we present implied-volatility forecasting as a controlled testbed where this gap is visible because the descriptors are pre-specified.