From Groups to Rings: Causal Evidence for Algebraic Decomposition in Grokked Transformers
Shurui Zheng ⋅ Fanhong Li ⋅ Zhixing Huang ⋅ ZIXI LI ⋅ Lei Ji
Abstract
Prior work has shown that grokked Transformers implement discrete Fourier transforms or group character representations for modular arithmetic. We unify these findings under the Wedderburn--Artin decomposition: using interchange intervention accuracy (IIA), a causal method stronger than ablation, we show that grokked models decompose computation along the Wedderburn components of the target algebra. Across 11 commutative algebras over $\F_p$ ($p = 2,3,5,7$)---including 6 with nontrivial Jacobson radical---all grokked models achieve raw IIA $\geq 0.97$. When models predict incorrectly, errors remain compartmentalized by component ($3.8\times$ chance), confirming that independence is structural, not a byproduct of accuracy. Training dynamics reveal that Wedderburn alignment emerges synchronously with grokking. For non-commutative groups, models exhibit \emph{selective} Wedderburn alignment; non-commutativity raises a capacity threshold that larger models overcome.
Chat is not available.
Successful Page Load