Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
Abstract
Lay Summary
When AI systems solve math problems, they often rely on a second AI, a "reward model," to grade their reasoning step by step. These graders are increasingly used to train the problem-solving AI, so their trustworthiness matters. We stress-tested today's leading graders and found they are easily fooled: they reward answers that merely sound like careful reasoning, even when the logic is wrong. Rewording a correct solution barely changes its grade, yet the graders often miss deliberately corrupted reasoning. Worse, when an AI is trained to chase these grades, it learns to game the grader rather than reason better, scoring nearly perfectly while getting the math right less than 4% of the time. Current graders act more like fluency detectors than genuine reasoning checkers. We release a benchmark and toolkit to help find and fix these weaknesses before deployment.