Linear Pin Representations and the Limits of Patching Evidence in a Searchless Chess Transformer
Abstract
An important question in mechanistic interpretability is whether neural networks trained on structured domains develop linear representations of abstract concepts, and whether intervening on these representations affects the model's output. We present a case study on the 270M searchless chess transformer of Ruoss et al. (2024), focusing on absolute pins. A linear probe identifies pin status well above bitboard, untrained-model, and random-label baselines. Mean-difference activation patching along the resulting pin direction reshapes the model's value-prediction output, but two controls reveal caveats to the behavioral interpretation. A placebo direction for the broader concept of geometric king–slider alignment produces a larger effect at matched magnitude despite being nearly orthogonal, and per-action scoring shows the model's recommended move rarely changes despite the much larger bucket-flip rate. Probe and patching peaks lie at different layers in both 270M and 9M variants. Beyond our chess-specific findings, our results provide a quantitative case study of how standard interpretability evidence types, probe accuracy, random-direction controls, and output-bucket flips, can each overstate behavioral causation.