A Two-Stage Cross-Attention Circuit for Spatial-Relation Binding in Small Diffusion Transformers
Juliana Li ⋅ Binxu Wang
Abstract
We mechanistically resolve how a small PixArt-mini DiT, in a controlled object-relation testbed, converts a textual spatial relation into image layout, identifying a two-stage cross-attention circuit at individual-head granularity. In this 6-layer, 36-cross-attention-head DiT (T5-XXL conditioning, $\sim 1000$ synthetic prompts), two early-layer heads $\{L_0H_0, L_1H_2\}$ implement spatial routing: their QK circuits convert each relation token into an $8\times 8$ directional attention pattern, and the corresponding OV writes carry that pattern into the residual stream. Zeroing $L_0H_0$ alone causes $+48$ pp damage to spatial accuracy ($0.84\to 0.35$); the pair causes $+64$ pp damage (super-additive $I=+0.086$, $95\%$ CI $[+0.015,+0.157]$). A four-method consensus search rejects every other tested head as a spatial-routing partner ($|I|\leq 1.3$ pp). A second object-binding stage is distributed across Layer 2: it has near-zero effect alone, but is strongly super-additive on color and shape when co-ablated with Layer 0 (a downstream stage that pair-ablation scored on spatial accuracy alone would miss). The circuit emerges via a rapid rise in projection magnitude during epochs 750-1000, with behavioral accuracy lagging by $\sim 200$ epochs. Within this setting, head-level causal circuit discovery extends to a small DiT, and scoring ablations on multiple metrics is needed to detect distributed downstream stages.
Chat is not available.
Successful Page Load