A Geometric Perspective on Stabilizing Value Conflict Resolution
Abstract
Large Language Models (LLMs) often struggle to navigate user requests that contain complex value conflicts. To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Geometrically, we show that CoT correlates with further smoothing the model's loss landscape in its sharpest direction, helping resolve the optimization instability of traditional scalar rewards. In terms of performance, we demonstrate via relevant downstream benchmarks that CoT may increase capability, demonstrating that value conflict-focused CoT has the potential to be an effective mechanism for better moral reasoning. To capitalize on this potential, we create a new value conflict-focused CoT design that further smooths the sharpest direction of the loss landscape and increases moral reasoning performance. This finding shows that explicitly modifying and improving the design of reasoning dynamics offers a promising avenue for advancing pluralistic alignment in LLMs.