Chiral Symmetry Breaking in Transformers: A Group-Equivariant Framework for Addressing the Reversal Curse via Adjoint Manifold Mappings
Abstract
Lay Summary
Large language models can often learn a fact in one direction but fail when asked the same fact in reverse. For example, a model may learn that one entity is related to another, yet struggle to identify the first entity when queried from the second. This paper studies why this happens and argues that the information may sometimes be present inside the model but difficult to access through ordinary text generation. We propose a method that makes the model’s internal representations easier to use for reverse factual queries by adding a structured readout mechanism. Experiments on inverse-relation benchmarks show that this approach improves access to reverse factual information in a controlled retrieval setting. The work suggests that some failures of language models may come not only from missing training data, but also from how models retrieve and express information they have already encoded.