Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
Abstract
Multilingual safety remains significantly imbalanced, leaving non-high-resource (NHR) languages vulnerable compared to robust high-resource (HR) ones. Moreover, the neural mechanisms driving safety alignment remain unclear despite observed cross-lingual representation transfer.In this paper, we find that LLMs contain a set of cross-lingual shared safety neurons (SS-Neurons), a remarkably small yet critical neuronal subset that jointly regulates safety behavior across languages. We first identify monolingual safety neurons (MS-Neurons) and validate their causal role in safety refusal behavior through targeted activation and suppression. Our cross-lingual analyses then identify SS-Neurons as the subset of MS-Neurons shared between HR and NHR languages, serving as a bridge to transfer safety capabilities from HR to NHR domains. We observe that suppressing these neurons causes concurrent safety drops across NHR languages, whereas reinforcing them improves cross-lingual defensive consistency. Building on these insights, we propose a simple neuron-oriented training strategy that targets SS-Neurons based on language resource distribution and model architecture. Experiments demonstrate that fine-tuning this tiny neuronal subset outperforms state-of-the-art methods, significantly enhancing NHR safety while maintaining the model's general capabilities.
Lay Summary
Why the internet isn't equally safe in all languages ? While popular languages like English have robust AI safety filters, many other languages remain vulnerable to harmful content. We still don't fully understand how AI models manage safety across different languages. In this study, we discovered a tiny but powerful group of "Safety Neurons" inside large language models that act like a universal safety bridge. These specific neurons are shared across both high-resource and low-resource languages. By identifying and reinforcing these "shared neurons," we developed a simple training strategy to protect under-represented languages. Our experiments show that upgrading just this tiny fraction of the model can significantly boost safety for all languages without damaging the AI’s intelligence or performance. This work helps ensure that AI safety is a right for everyone, regardless of the language they speak.