HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation
Abstract
Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathematical rigor; (2) expressiveness to specify objectives and constraints, and (3) the ability to define tactics to achieve objectives. However, these benefits remain largely unexplored in the context of Multi-Agent Reinforcement Learning (MARL). This paper introduces HyPOLE, a novel framework for MARL under partial observability, where learning is guided by the expressive power of the so-called hyperproperties and, in particular, the temporal logic HyperLTL. We integrate Centralized Training for Decentralized Execution (CTDE) techniques with HyPOLE to synthesize decentralized policies, and our evaluation on SMAC, MessySMAC, and WildFire benchmark demonstrates clear advantages over baselines.
Lay Summary
Many important tasks require a team of agents, such as robots, drones, or autonomous vehicles, to work together. However, each agent usually sees only a small part of the environment, which makes it difficult to coordinate and make good decisions. In this paper, we introduce HyPOLE, a method that helps guide these agents using mathematical specifications that describe what they should do and what they should avoid. Instead of relying only on trial and error, HyPOLE uses these specifications to help agents learn safer and more coordinated behavior. This can be useful in settings where teams of intelligent systems need to work together, follow rules, and reach their goals even when no single agent has the full picture.