Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
Abstract
Lay Summary
“Should I prioritize my family or my personal goals?” People from different cultures may answer this question differently, revealing the values that shape how they think and communicate. As LLMs are used worldwide, it is important to know whether they appropriately reflect such cultural nuances. Yet most evaluations still rely on rigid surveys, asking models to rate statements like “How important is family to you on a scale of 1 to 5?” These surveys miss how people usually interact with AI: through open-ended writing, where values emerge indirectly. We introduce DOVE, a framework for evaluating cultural alignment in open-ended writing rather than fixed-choice answers. DOVE identifies culturally meaningful value codes, such as independence, duty, or personal ambition, from real-world texts and compares how often those value codes appear in human-written texts and AI-generated responses. Across 12 LLMs and four cultures, DOVE predicted culturally related model behavior more accurately than existing methods and remained reliable with relatively small datasets.