Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
Abstract
Lay Summary
(1) Large reasoning models can solve complex tasks, but they struggle to recognize the boundaries of their own knowledge and uncertainty. We asked whether a model could generate useful learning signals not only from final answers, but also from predicting how it succeed. (2) To study this, we designed a parallel reasoning pipeline where the model simultaneously answers a question and predicts meta-information about its own reasoning process, including estimated solving success rate, concept usage keywords, and expected solution length. These self-assessments form an auxiliary reward pathway that operates alongside standard answer matching. (3) Our findings suggest that meta-prediction can improve model performance even without additional supervision beyond question-answer pairs. By learning to estimate its own reasoning quality and knowledge boundaries, the model gains a richer training signal that complements direct correctness rewards. This opens a path toward reasoning systems that are not only stronger problem solvers, but also better calibrated and more self-aware during learning.