Skip to yearly menu bar Skip to main content


Thinking Reward Models Have a Self-Blind Spot: Chain-of-Thought Reward Models Cannot Detect Their Own Manipulation

Oliver Y Chen

Abstract

Chat is not available.