Look Before You Steer: Geometry Predicts SAE Feature Steerability
Muhammad Khan ⋅ Shlok Channawar ⋅ Akshaj Gurugubelli ⋅ Girish Gupta ⋅ Aditya Shah
Abstract
Steering with SAE features requires per-feature coefficient tuning, which currently demands intervention sweeps. We ask whether properties of the SAE itself, computable before any forward pass, predict which features will be cheap or expensive to steer. We show that variation in SAE feature steerability is partially predicted by decoder-space geometry: neighbor density and maximum cosine similarity to nearby decoder directions --- both computable from the SAE weight matrix before any intervention --- rank features by how much steering they require for a fixed behavioral effect ($\rho$ up to $-0.546$, $p < 10^{-6}$, AUROC $0.610$--$0.822$ across conditions; the signal is rank-based, consistent with grid discreteness). This geometry--steerability relationship replicates across two Gemma-2 model scales (2B and 9B), two SAE widths (16K and 65K), and is detectable cross-architecturally on Llama-3.1-8B-Instruct ($\rho = -0.266$, $n = 300$). On Qwen3-8B with BatchTopK SAEs, geometry predicts whether a feature is steerable at all but not the continuous ordering among responsive features, revealing a boundary condition tied to SAE training regime. The signal weakens at deep proportional layer depth in both models, where the cost of steering exceeds our intervention budget --- a consistent depth boundary. These results provide preliminary evidence that pre-steering geometry can partially inform coefficient selection, offering a path toward screening features for controllability before deployment.
Chat is not available.
Successful Page Load