In-Context Learning Amplifies a Latent Symbolic Circuit
Melissa Wessel
Abstract
Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to $8\times$ from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1\% to 56\% at 0-shot and 17\% to 88\% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86\%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.
Chat is not available.
Successful Page Load