Behavior Is Not Representation: Hypothesis Tests for Concept Steering in LLMs
Siddharth Shukla
Abstract
We formulate the $Linear Representation Hypothesis$ (LRH) as a falsifiable hypothesis test for LLM steering: if steering vectors encode semantic concepts, their Sparse Autoencoder (SAE) decompositions should recover the corresponding concept-level SAE features at rates significantly above chance. Across 950 concepts from AxBench and the security-domain CryptoBench, four steering methods, and a 5-layer landmark sweep of Gemma-2-2B-IT, we find systematic failure of feature recovery (Top-20 recovery $\leq 0.4\%$; binomial $p > 0.05$ after Bonferroni correction across all 20 method $\times$ layer configurations). Analytical methods exhibit apparent structural alignment, but this alignment is dominated by formatting and syntax features rather than semantic concepts. Gradient-trained steering vectors produce the largest behavioral effects while remaining nearly orthogonal to concept-specific SAE features. We introduce $Structural Specificity Testing$ (SST), a protocol for validating whether steering vectors recover concept-specific representations beyond calibrated null baselines. These results reveal a structural-behavioral disconnect in current steering evaluations and motivate hypothesis-testing protocols that jointly validate behavioral efficacy and structural specificity.
Chat is not available.
Successful Page Load