Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
Abstract
Lay Summary
As AI assistants become widely used, they are increasingly expected to reflect human values (e.g., caring for others, striving for achievement, respecting tradition...). We can make models express these values using two different methods: either by training them to answer that way, or by inserting an explicit prompt prompt to express a value. However, while these two methods are often treated as to be interchangeable, it is not really clear whether they work in the same way inside the model. We study this question by looking at the internal activity of language models while they express different human values. We find that the two types of value expression partly share a common mechanism: both capture meaningful value concepts and relationships among values. However, they also differ in important ways. Intrinsic values are associated with mechanisms that support more varied and natural responses, whereas prompted values are more strongly linked to mechanisms that make the model follow instructions. This distinction matters because training and prompting are two common ways to shape AI behavior. Our results suggest that they are not simple substitutes: both have their own strengths, and understanding how these two mechanisms are similar and different can help people choose the right approach for different situations.