VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
Abstract
Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path, yet three gaps persist: extraction often ignores hierarchical structure, evaluation detects presence but not calibrated intensity, and steerability of LLMs at controlled intensities remains insufficiently understood. To address these limitations, we introduce VALUEFLOW, a unified framework that spans extraction, evaluation, and steering with calibrated intensity control. The framework integrates three components: (i) HiVES, a hierarchical value embedding space that captures intra- and cross-theory value structure; (ii) the Value Intensity DataBase (VIDB), a large-scale resource of value-labeled texts with intensity estimates derived from ranking-based aggregation; and (iii) an anchor-based evaluator that produces consistent intensity scores for model outputs by ranking them against VIDB panels. Using VALUEFLOW, we conduct a comprehensive large-scale study across ten models and four value theories, identifying asymmetries in steerability and composition laws for multi-value control. This paper establishes a scalable infrastructure for evaluating and controlling value intensity, advancing pluralistic alignment of LLMs.
Lay Summary
Large language models are used by many people with different beliefs, priorities, and cultural backgrounds. But today’s methods for making AI behave in helpful ways often rely on simple preferences, such as which answer a person likes better, and may miss the deeper values behind those choices. In this work, we study how language models can better recognize and respond to human values. We introduce VALUEFLOW, a framework that helps researchers see what values appear in text, measure how strongly those values are expressed, and guide a model to reflect a chosen value more or less strongly. For example, instead of only asking a model to be “fair” or “helpful,” VALUEFLOW can test whether the model expresses that value weakly, strongly, or in balance with other values. Our approach gives the model a clearer map of different human values, collects many real examples showing those values at different strengths, and compares new model answers with these examples. Our results show that different models vary greatly in how easily they can be guided, and that some values are much harder to change than others. This can help researchers build AI systems that respond to diverse human values in a more stable, transparent, and understandable way.