Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Abstract
Lay Summary
While large language models have become better at solving complex problems through trial-and-error learning, current methods often struggle when tasks move beyond simple math. Most existing systems only reward a model if the final answer is perfectly correct, which limits the model’s ability to explore different ways of thinking and solve problems in diverse fields like chemistry or physics. To bridge this gap, we developed RGR-GRPO, a new training framework that uses "rubrics"—detailed sets of scoring criteria—to guide the model. Instead of a simple "right or wrong" grade, our system provides fine-grained feedback throughout the reasoning process, acting like a teacher who gives partial credit and helpful hints. Our experiments across 14 different subject areas show that this approach significantly boosts performance, with improvements of up to 8.4% in complex science tasks. This research makes AI reasoning more robust and versatile, moving us closer to AI assistants that can tackle a wide range of sophisticated real-world challenges across multiple academic and professional domains.