Cross-Model Circuit Discovery
Harrish Thasarathan ⋅ Matthew Kowal ⋅ Thomas Fel ⋅ Konstantinos Derpanis
Abstract
Consider two large vision models. They process the same image and both correctly predict the class ``rabbit.'' How much of the circuit computation along the way was shared? Model diffing offers a natural lens on this question. So far, however, it has largely operated on a single layer and at the level of representations rather than circuits. In this work, we introduce Universal Circuits (UCs), enabling model diffing at the circuit level. Specifically, we extend CLTs across both layers and models, with losses that encourage sparsity for interpretability and output fidelity for faithfulness. We train UCs between pairs of standard large vision models. We find that a compact cross-model intersection of pruned class circuits, typically a few hundred universal (shared) features per pair, produces $89$-$98\%$ of full-circuit classification accuracy in both models, while a complementary set of universal features (hundreds per pair) is kept by only one model's circuit, reflecting how each model weights shared concepts differently in its own representations. To demonstrate a downstream use of UC, we perform model \textit{surgery}: a class is successfully transferred from one model to another with no gradient steps taken in the new model. More broadly, our work demonstrates that large vision models leverage shared multi-layer algorithms for downstream performance, and that these algorithms can be discovered, compared, and reused across models.
Chat is not available.
Successful Page Load