Reinforcement Learning via Self-Distillation
Abstract
Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with verifiable rewards (RLVR) learn only from a scalar outcome reward per attempt, creating a severe credit-assignment bottleneck. Many verifiable environments actually provide rich textual feedback, such as runtime errors or judge evaluations, that explain why an attempt failed. We formalize this setting as reinforcement learning with rich feedback and introduce Self-Distillation Policy Optimization (SDPO), which converts tokenized feedback into a dense learning signal without any external teacher or explicit reward model. SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy. In this way, SDPO leverages the model's ability to retrospectively identify its own mistakes in-context. Across scientific reasoning, tool use, and competitive programming on LiveCodeBench v6, SDPO improves sample efficiency and final accuracy over strong RLVR baselines. Notably, SDPO also outperforms baselines in standard RLVR environments that only return scalar feedback by using successful rollouts as implicit feedback for failed attempts. Finally, applying SDPO to individual questions at test time accelerates discovery on difficult binary-reward tasks, achieving the same discovery probability as best-of-k sampling or multi-turn conversations with 3x fewer attempts.
Lay Summary
Large language models are often improved by letting them try a task, receive a score, and update from that score. But for tasks like coding, math, or scientific reasoning, a simple “right” or “wrong” signal can hide useful information about why the model failed. For example, a programming system may return a runtime error or a failed test case, which can point directly to the mistake. We introduce Self-Distillation Policy Optimization, a training method that helps a model learn from this richer feedback without needing a stronger external teacher. The key idea is to let the model look back at its own answer after seeing the feedback, then use its feedback-informed judgment to teach its original version which parts of the answer were more or less promising. This turns sparse feedback into more detailed guidance during training. Across scientific reasoning, tool use, and competitive programming tasks, this approach learns more efficiently than strong reinforcement-learning baselines. It can also help models solve difficult problems with fewer attempts, making training and test-time problem solving more effective.