Cross-Lingual Emergent Misalignment: Shared Multilingual Circuits Can Propagate Safety Failures
Abstract
Modern large language models exhibit language-specific processing components alongside shared, language-independent abstract representations of concepts. As multilingual LLMs are deployed globally, it is critical to understand whether misaligned behaviour can propagate across languages through shared pathways. Recent work has shown that extremely small, low-rank fine-tuning on harmful datasets can induce emergent misalignment (EM) in a mechanistically structured and a robust way. In this work, we fine-tune the multilingual Tiny Aya model family (3.35B parameters) on insecure-text datasets in English and evaluate EM transfer across eight typologically diverse languages spanning distinct regions, scripts, and resource tiers: English, Portuguese, Turkish, Hindi, Marathi, Urdu, Hausa, and Yoruba. We find that EM transfer is present in a structured and non-uniform manner across all languages, that there appears to be a shared misalignment direction inside the network, and that the strongest EM transfer is observed for languages whose internal representations overlap more strongly with this direction.