What Structural Inductive Bias Helps Transformers Reason Over Knowledge Graphs?
Jonas Petersen ⋅ Camilla Mazzoleni ⋅ Gian-Alessandro Lombardi ⋅ Federico Martelli ⋅ Riccardo Maggioni
Abstract
What structural inductive bias helps transformers reason over knowledge graphs? Through controlled ablations of a minimal transformer modification with four independently removable components (sparse adjacency masking, edge-type biases, query scaling, value gating), we isolate which structural signals drive multi-hop reasoning. Our finding is sharp: sparse adjacency masking alone accounts for the dominant share of improvement over unmasked transformers ($+72.5$pp on 3-hop MetaQA, $+45.5$pp on WebQSP, $+53.9$pp on CWQ), while learned relation parameters add only modest refinement and can actively hurt without structural guidance. A zero-shot experiment provides architecturally independent corroboration: masking-based attention degrades $4.0\times$ less than relation-specific weights when edge types are held out. The useful inductive bias for multi-hop KGQA is predominantly topological, not relational.
Chat is not available.
Successful Page Load