Harmonized Dual Policy Improvement for Modelic Reinforcement Learning
Abstract
Lay Summary
Modern AI systems that learn to make decisions, such as robots or autonomous agents, often rely on two complementary abilities: quickly reacting based on past experience, and planning ahead by imagining possible future outcomes. Recent research has shown that combining these two abilities can greatly improve performance. However, in practice, they often give conflicting learning signals, which can destabilize training and prevent the system from reaching its full potential. In this work, we study why this conflict happens and propose a new method to better coordinate the two learning processes. Our approach carefully balances short-term decision making with long-term planning so that they work together instead of interfering with each other. We also provide theoretical analysis showing that both learning strategies are ultimately trying to achieve the same goal. We evaluate our method on a range of challenging simulated control tasks, including humanoid and quadruped locomotion. The results show that our approach makes training more stable and significantly improves final performance compared with existing methods. These findings provide a more reliable way for future AI systems and robots to learn complex behaviors efficiently.