All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs
Abstract
In this paper, we present empirical and theoretical evidence against a central but largely implicit assumption in circuit and sheaf discovery (CSD), which we term the Functional Anisotropy Hypothesis: the idea that functions in large language models (LLMs) are localised to a unique or near-unique internal mechanism. We show that a single LLM task can instead be supported by multiple, structurally distinct circuits or sheaves that are simultaneously faithful, sparse, and complete. To systematically uncover such competing mechanisms, we introduce Overlap-Aware Sheaf Repulsion, a method that augments the CSD objective with an explicit penalty on structural overlap across multiple discovery runs, enabling the discovery of circuits or sheaves with strong task performance but minimal shared structure across a plethora of common CSD benchmarks. We find that this phenomenon becomes increasingly pronounced as the number of discovered sheaves grows and persists robustly across major CSD methods. We further identify an ultra-sparse three-edge sheaf and show that none of its edges is individually indispensable, undermining even weakened notions of canonical or essential components. To explain these findings, we propose a Distributive Dense Circuit Hypothesis and provide a theoretical analysis demonstrating that non-unique, low-overlap circuit explanations arise naturally from high-dimensional superposition under mild assumptions. Together, our results suggest that mechanistic explanations in LLMs are inherently non-canonical and call for a rethinking of how CSD results should be interpreted and evaluated.
Lay Summary
Large language models are often interpreted as using a specific internal "circuit" or "sheaf" to perform a task. Our paper challenges this assumption. We show that the same task can often be carried out by many very different internal mechanisms inside the same model, even when those mechanisms barely overlap structurally. To study this systematically, we introduce a method called Overlap-Aware Sheaf Repulsion (OASR) that augments DiscoGP, which explicitly searches for alternative low-overlap circuits or sheaves that still solve the task well. Across multiple benchmarks, by exploiting instabilities and arbitrary choices in existing circuit discovery methods, we consistently uncover many competing explanations rather than a single canonical one. We also identify an extremely small three-edge mechanism for a standard language-model task, but show that even its components are not uniquely essential. Altogether, our results suggest that computation in large language models is more distributed and non-unique than current mechanistic interpretability assumptions imply. We further provide a theoretical framework explaining why multiple low-overlap explanations can naturally arise in large language models.