EvoMAS: Evolutionary Generation of Multi-Agent Systems
Abstract
Large language model (LLM)-based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures remains labor-intensive, brittle, and hard to generalize. Existing automatic MAS generation methods either rely on code generation, which often leads to executability and robustness failures, or impose rigid architectural templates that limit expressiveness and adaptability. We propose Evolutionary Generation of Multi-Agent Systems (EvoMAS), which formulates MAS generation as structured configuration generation. EvoMAS performs evolutionary generation in configuration space. Specifically, EvoMAS selects initial configurations from a pool, applies feedback-conditioned mutation and crossover guided by execution traces, and iteratively refines both the candidate pool and an experience memory. We evaluate EvoMAS on diverse benchmarks, including BBEH, SWE-Bench, and WorkBench, covering reasoning, software engineering, and tool-use tasks. EvoMAS consistently improves task performance over both human-designed MAS and prior automatic MAS generation methods, while producing generated systems with higher executability and runtime robustness. EvoMAS outperforms the agent evolution method EvoAgent by +10.5 points on BBEH reasoning and +7.1 points on WorkBench. With Claude-4.5-Sonnet, EvoMAS also reaches 79.1% on SWE-Bench-Verified, matching the top of the leaderboard. Code is available at https://github.com/amazon-science/EvoMAS
Lay Summary
AI systems that coordinate multiple specialized agents (where one agent writes code, another reviews it, and a third tests it) can solve complex problems better than any single agent alone. But designing these multi-agent teams currently requires significant human expertise: deciding which agents to include, what roles they play, and how they communicate. We developed EvoMAS, a system that automatically designs these agent teams through a process inspired by biological evolution. Starting from a small set of human-designed team templates, EvoMAS iteratively improves them: it tests different configurations, identifies what works, and combines successful elements, much like how natural selection produces organisms well-adapted to their environments. Crucially, it operates on structured blueprints rather than raw code, which makes the process reliable and the resulting teams interpretable. Across tasks spanning mathematical reasoning, software engineering, and workplace automation, EvoMAS discovers agent teams that outperform both human-designed systems and prior automated approaches. This suggests that evolutionary search over team structures is a promising path toward AI systems that can automatically organize themselves to tackle complex real-world problems.