Predictable Steering: Leveraging Geometric Proxies for Layer Selection and Per-Instance Success Prediction
Nishkal Hundia ⋅ Swastik Agrawal ⋅ Navita Goyal ⋅ Sarah Wiegreffe
Abstract
Representation steering methods like Contrastive Activation Addition (CAA) offer a lightweight approach to controlling LLM behavior, but their effectiveness varies widely across concepts, datasets, and layers, making principled deployment costly without exhaustive hyperparameter sweeps. We build on the discriminability index $d'$, a compute-cheap geometric measure of how well constrastive training activations separate along the steering direction, and demonstrate uses beyond its established role as a dataset-level steerability predictor. First, we show that $d'$ reliably identifies the optimal steering layer for a given behavior across five multiple-choice behavioral question datasets on Gemma-2-9B-IT. Second, we show that for high-$d'$ datasets, the position of a generated token's representation on the difference-of-means line predicts steering success on that individual example, with per-layer MCC tracking $d'$ closely across layers. Together, these findings suggest that $d'$, provides actionable guidance for both layer selection and per-example assessment of steering success without requiring ground-truth labels or exhaustive downstream evaluation. Code: https://tinyurl.com/2tx8y6yu
Chat is not available.
Successful Page Load