Invited Talk: LLMs are aligned - but with whom?, by Ms. Pardis Sadat Zahraei (UIUC)
Abstract
Safety alignment in large language models is typically framed as a technical problem with technical solutions. In this talk, we argue it is also a cultural one. Drawing on our own work and a growing body of evidence on cross-cultural alignment failures, we examine what happens when alignment training meets a pluralistic, multi-value world. We introduce the concept of the alignment veto, where a model holds accurate cultural knowledge internally but a safety-trained gate prevents its expression, and discuss how this differs from representational bias failures, which require fundamentally different interventions. We also challenge the common assumption that native-language prompting improves cultural accuracy, and ask broader questions: whose values should LLMs represent, who gets to decide, and what would a genuinely community-situated alignment framework look like?