Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
Abstract
Lay Summary
(1) Problem To enable multiple AIs, such as autonomous vehicle fleets or search-and-rescue robots, to collaborate seamlessly, they require extremely flexible "brains." While diffusion models, the technology behind AI image generation, are remarkably powerful and can provide the necessary flexibility, their underlying complex mathematical mechanisms cause AIs to easily get "stuck" when exploring new strategies. For this reason, this powerful technology has previously been difficult to apply to real-time guidance of AI team collaboration. (2) Solution To untangle this mathematical deadlock, we developed a new framework called OMAD, the first successful method to apply diffusion models to online training of AI teams. We cleverly provided the AIs with a set of "relaxed requirement" training rules, allowing them to explore boldly and maintain coordination even when they cannot compute exact mathematical probabilities. Under this framework, AIs receive centralized guidance during training as a collective, yet each AI can independently make smart and coordinated decisions during task execution. (3) Impact This breakthrough delivers remarkable results. Across ten complex team task benchmarks, OMAD outperformed all existing methods, setting entirely new state-of-the-art records. More importantly, it dramatically accelerated AI team learning speed by 2.5 to 5 times. This research unlocks the potential of generative AI in multi-agent systems, paving the way for more efficient robot team deployments in the real world.