Behavior Steering via Layer-to-Layer Jacobian Singular Vectors
Omar B Ayyub ⋅ John Strand
Abstract
The map of how activations at one source layer in an LLM impact activations at a later target layer, the layer-to-layer Jacobian, yields cheap steering vectors via its top right singular vectors. Block power iteration recovers the top-$k$ such vectors in roughly 15 forward passes per source/target pair, giving this method its name: Power Steering. We benchmark Power Steering against Contrastive Activation Addition (CAA) and the nonlinear layer-to-layer method MELBO on Qwen3-14B across seven Anthropic advanced-AI-risk evaluations, under both first-token logit-difference and LLM-judged sampled-generation metrics. The two layer-to-layer methods (Power Steering and MELBO) produce stronger in-class steering and cross-evaluation transfer than CAA. Power Steering closely tracks MELBO under logit metrics, with a modest advantage to MELBO under sampled generation, at a fraction of the per-pair cost. This cheap per-pair cost lets us map every source/target pair in the model: from a single phishing-email prompt, the resulting model-map surfaces anti-refusal vectors that generalize to a subset of AdvBench harm categories. Steering is most easily found on prompts with decision forks but can also surface latent behaviors.
Chat is not available.
Successful Page Load