Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Abstract
A reliable reward model is essential for aligning large language models (LLMs) with human preferences through reinforcement learning from human feedback (RLHF). However, standard reward models are susceptible to spurious features that are not causally related to human labels. This can lead to reward hacking, where high predicted reward does not translate into better behavior. In this work, we address this problem from a causal perspective by proposing a factored representation learning framework that decomposes the model’s contextual embedding into (1) causal factors that are sufficient for reward prediction and (2) non-causal factors that capture reward-irrelevant attributes such as length or sycophantic bias. The reward head is then constrained to depend only on the causal component. In addition, we introduce an adversarial head trained to predict reward from the non-causal factors, while applying gradient reversal to discourage them from encoding reward-relevant information. Experiments on both mathematical and dialogue tasks demonstrate that our method learns more robust reward models and consistently improves downstream RLHF performance over state-of-the-art baselines. Analyses on length and sycophantic bias further validate the effectiveness of our method in mitigating reward hacking behaviors.
Lay Summary
Large language models are often improved using human feedback, where another model learns to score which answers people prefer. However, this scoring model can sometimes be fooled by shortcuts. For example, it may give higher scores to answers that are longer or that simply agree with the user, even when those answers are not actually better. When the language model is trained to maximize such scores, it may learn to game the scoring system rather than become more helpful or accurate. In this work, we propose CausalRM, a method that helps the scoring model focus on the real reasons an answer should be preferred, while reducing the influence of distracting patterns such as length or overly agreeable wording. The key idea is to separate useful information for judging answer quality from information that should not affect the score. We test our method on mathematical reasoning and open-ended dialogue tasks. The results show that CausalRM makes the scoring model more reliable, improves the final language model trained with feedback, and reduces common shortcut-seeking behaviors.