G$^2$RPO: Geometric GRPO; Escaping LLM's Reasoning Rut to Break Accuracy--Entropy Trade-off
Abstract
Lay Summary
When we train language models to solve problems with automatically checkable answers, they often get better by becoming narrower. They may discover one reliable way to solve a problem and repeat it, even when several correct approaches exist. We call this a reasoning rut. This paper explains why that rut appears in GRPO, a popular reinforcement-learning method for reasoning models: its updates naturally push correct solution strategies to compete until one dominates. We introduce G²RPO, a small change to the training signal that gives extra credit to correct approaches the model is currently underusing. A balancing term keeps the model’s pressure to abandon wrong answers close to the original training dynamics, so diversity is not bought by sacrificing accuracy.