Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy Optimization
Abstract
To improve generalization and resilience in human–robot collaboration (HRC), robots must contend with diverse combinations of human behaviors and contexts, motivating multi-agent reinforcement learning (MARL). However, inherent heterogeneity between robots and humans creates a rationality gap (RG), where decentralized policy updates deviate from cooperative joint optimization. The resulting learning problem is a general-sum differentiable game, so independent policy-gradient updates can oscillate or diverge without added structure. We propose heterogeneous-agent Lyapunov policy optimization (HALO), a framework that stabilizes decentralized MARL by enforcing Lyapunov-based contraction in policy-parameter space. Unlike Lyapunov-based safe RL, which targets state/trajectory constraints in constrained Markov decision processes, HALO uses Lyapunov certification to stabilize decentralized policy learning. HALO rectifies decentralized gradients via optimal quadratic projections, ensuring monotonic contraction of RG and enabling effective exploration of open-ended interaction spaces. Extensive simulations and real-world humanoid-robot experiments show that this certified stability improves generalization and robustness in collaborative corner cases.
Lay Summary
Many real-world jobs require humans and robots to work together, especially in unpredictable places like homes, hospitals, factories, and public spaces. The challenge is that people and robots behave and learn differently. Humans can act in unexpected ways, and this mismatch can make robot learning unstable, causing robots to overreact, fail to adapt, or make teamwork worse. We propose HALO, a new learning method that helps robots collaborate with human partners whose behavior is only partly predictable. HALO keeps the robot’s learning process stable while steadily reducing disagreement between the human and robot, helping the robot become a more reliable teammate over time. We tested and validated HALO in simulation and real-world humanoid robot experiments across multiple tasks. HALO improves reliability, especially when people behave unexpectedly, supporting safer human-robot collaboration in high-impact settings.