Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment
Abstract
Lay Summary
AI models are trained to abstain from “unsafe” responses and prefer “safe” ones: for example, not generating violent images in response to innocent prompts. But who decides what is “safe”? Researchers collect preferences from a large number of people and aggregate their vote to determine whether something should be considered safe or not. However, they rarely describe who those people are: we found only 8 datasets that report both demographic and geographic information about the raters. Previous research showed that factors like age, gender, and ethnicity all affect safety ratings. Our paper further shows that the rater’s cultural background (like where someone was born, where they live, or their nationality) shapes safety judgments in ways the other demographics don't capture. In fact, roughly 10% of items in the datasets we looked at were “culturally sensitive”: flagged as unsafe by only one cultural group. This means that without cultural diversity in the people training AI, models risk being unsafe for entire populations whose cultural perspectives weren't included. We tested whether AI could simply simulate these diverse perspectives and found the models unreliable. However, we proposed an AI model that helps flag the items where cultural groups are most likely to disagree, so those can be prioritized for human review.