Circuit Modularity Predicts Compositional Generalization: Theory and Evidence from Transformers
Kaustubh Bukkapatnam ⋅ Siddharth Karuturi
Abstract
We ask whether the internal circuit structure of a neural network predicts its ability to generalize compositionally to novel combinations of seen primitives. We introduce the Compositional Modularity Index (CMI), a scalar measure of how consistently a network's intermediate representations respond to changes in individual components of a compositional input, estimated via causal activation patching. We prove a generalization bound showing that a model with $CMI \ge \mathcal{M_0}$ on training data incurs a compositional out-of-distribution (OOD) generalization gap of at most $O(1 - \mathcal{M}_{0})$ over the i.i.d. gap, plus standard statistical terms. This bound is tight to a constant factor. Empirically, we measure CMI for four transformer architectures---a standard Transformer (TX), a Universal Transformer (UT), a Relation Network enhanced Transformer (RELNET), and a modular variant (MODTX)---across the SCAN and COGS compositional generalization benchmarks. CMI correlates strongly with OOD accuracy ($r = 0.94$, $p < 10^{-6}$), predicts which layers drive compositional behavior, and rises earlier during training for architectures with stronger inductive biases toward modularity. Our results suggest CMI is a useful diagnostic for compositional generalization without requiring OOD labels.
Chat is not available.
Successful Page Load