Layer-Wise Category Structure in Large Language Models : A Cross-Architecture Analysis of Feature Decodability
Guus Bouwens
Abstract
We study where category-relevant information becomes linearly decodable across layers in four instruction-tuned language models. Using residual-stream activations from 215 prompts spanning 16 task categories, we train 128 layer-wise linear probes and compare coarse early, middle, and late processing trends across architectures. We find a consistent broad representational organization in which some categories (for example, spatial navigation and logical reasoning) become decodable earlier than others, alongside meaningful architecture-specific differences in late-layer behavior such as Mistral-7B's late accuracy drop and Llama-8B's stronger confidence concentration. To keep the paper at a descriptive level, we emphasize phase-bins rather than exact layer rankings: with 16 categories, the current design has 80\% power only for cross-model rank correlations of about $\rho \approx 0.65$, and fine-grained orderings are sensitive to preprocessing choices. In particular, rank agreement with the default z-score pipeline drops to $\rho=0.287$ with raw activations and $\rho=0.101$ with per-sample L2 normalization. Sparse autoencoder analyses provide a complementary unsupervised-style cross-check of category-selective structure, but we do not claim a localized causal mechanism. The resulting contribution is a calibrated map of broad layer-wise category structure, together with explicit boundaries on what can and cannot be concluded from the current four-model, 215-prompt dataset.
Chat is not available.
Successful Page Load