Does Your Reasoning Model Implicitly Know When to Stop Thinking?
Abstract
Recent advancements in large reasoning models (LRMs) have greatly improved their capabilities on complex reasoning tasks through Long Chains of Thought (CoTs). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. Recent studies show that longer reasoning chains are frequently uncorrelated with correctness and can even be detrimental to accuracy. In a further in-depth analysis of this phenomenon, we surprisingly uncover and empirically verify that LRMs implicitly know the appropriate time to stop thinking, while this capability is obscured by current sampling paradigms. Motivated by this, we introduce SAGE (Self-Aware Guided Efficient Reasoning), a novel sampling paradigm that unleashes this efficient reasoning potential. Furthermore, integrating SAGE as mixed sampling into group-based reinforcement learning (SAGE-RL) enables SAGE-RL to effectively incorporate SAGE-discovered efficient reasoning patterns into standard pass@1 inference, markedly enhancing both the reasoning accuracy and efficiency of LRMs across multiple challenging mathematical benchmarks.
Lay Summary
Recent AI models have become very good at solving complex reasoning problems by thinking through long, step-by-step processes. However, these long thinking chains often include unnecessary steps, slowing down computation and sometimes even reducing accuracy. In our research, we discovered that these models actually “know” when they have thought enough to reach a correct answer, but this ability is hidden by the way current methods generate their reasoning. To address this, we introduce SAGE (Self-Aware Guided Efficient Reasoning), a new method that helps models stop thinking at the right time, avoiding wasted steps. When combined with reinforcement learning in SAGE-RL, models can learn to apply these efficient reasoning patterns during standard problem-solving. Across multiple challenging math benchmarks, this approach improves both the speed and correctness of AI reasoning, making models smarter and faster without unnecessary computation.