CAMEL: Confidence-Gated Reflection for Reward Modeling
Abstract
Reward models play a fundamental role in aligning large language models with human preferences. Existing methods predominantly follow two paradigms: scalar discriminative preference models, which are efficient but lack interpretability, and generative judging models, which offer richer reasoning at the cost of higher computational overhead. We observe that the log-probability margin between verdict tokens strongly correlates with prediction correctness, providing a reliable proxy for instance difficulty without additional inference cost. Building on this insight, we propose CAMEL, a confidence-gated reflection framework that performs a lightweight single-token preference decision first and selectively invokes reflection only for low-confidence instances. To induce effective self-correction, we train the model via reinforcement learning with counterfactual prefix augmentation, which exposes the model to diverse initial verdicts and encourages genuine revision. Empirically, CAMEL achieves state-of-the-art performance on three widely used reward-model benchmarks with 82.9% average accuracy, surpassing the best prior model by 3.2% and outperforming 70B-parameter models using only 14B parameters, while establishing a strictly better accuracy-efficiency Pareto frontier.
Lay Summary
AI assistants are often improved by teaching them which answers people prefer. This paper studies how to make that preference judgment both accurate and efficient. Existing methods either make fast but opaque decisions, or generate longer explanations that are more costly. We introduce CAMEL, a method that first makes a quick choice between two answers and then checks how confident it is. When the choice is easy, CAMEL stops immediately. When the choice is uncertain, it takes a second look before making a final decision. This selective reflection helps the model spend extra effort only where it is useful. Experiments show that CAMEL improves judgment accuracy over strong existing methods while using less computation.