The Unembedding Bottleneck: A Mechanistic Analysis of Single-Digit Counting in LLMs
Satwik Sunnam ⋅ Raghav Magazine ⋅ Vatsalya Singh ⋅ Lavanya Kotha ⋅ Xingjian Li ⋅ Min Xu
Abstract
Large language models consistently fail at elementary counting tasks despite strong performance on complex reasoning benchmarks. Recent work has characterized counting failures behaviorally or identified circuits through which counting succeeds, but a mechanistic account of where and why counting breaks down remains absent. In this work, we provide such an account by tracing the count signal from embedding to output and identifying the geometric bottleneck that prevents correct readout for single digit counting problems. Using a dual-metric framework combining Signal-to-Noise Ratio (SNR) with linear probing, we show that counting information in models is computed correctly and persists in a linearly decodable subspace through the final layer. The failure in counting is not representational but geometric, i.e., the digit-token columns of the unembedding matrix $W_U$ are orthogonal to the count subspace in the residual stream, rendering correctly computed counts unreadable at the output. We term this as the $\textbf{Unembedding Bottleneck}$ and to causally validate it, we apply a $\textbf{Probe-Informed LoRA}$ correction targeting $W_U$, which restores single-digit counting accuracy by up to 80\% across six models while updating only ${\sim}0.001$\% of parameters and preserving general capabilities. Our findings reveal a broader class of failure in LLMs where information is internally present but geometrically inaccessible to the output projection.
Chat is not available.
Successful Page Load