CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
Abstract
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave less cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings. Indeed, our experiments show that recent models---with or without reasoning enabled---consistently defect in single-shot social dilemmas. To tackle this safety concern, we present the first comparative study of game-theoretic mechanisms designed to enable cooperative outcomes between rational agents in equilibrium. Across four social dilemmas testing distinct components of robust cooperation, we evaluate four families of mechanisms: (1) repeating the game for many rounds, (2) reputation systems, (3) third-party mediators to delegate decision making to, and (4) contract agreements for outcome-conditional payments between players. Among our findings, we establish that contracting and mediation are most effective in achieving cooperative outcomes between capable LLM models, and that repetition-induced cooperation deteriorates drastically when co-players vary. Moreover, we demonstrate that the mechanisms become more effective under evolutionary pressures to maximize individual payoffs.
Lay Summary
As artificial intelligence (AI) agents become more advanced and handle real-world tasks together, a major safety concern has emerged: smarter AI models tend to behave more selfishly. When placed in situations where cooperation benefits everyone but there are alternative actions that reward the individual, today's AI models consistently choose to look out only for themselves, leading to collective failure. To tackle this, we conducted the first comprehensive study to test how different interaction designs can guide AI agents toward cooperation. We evaluated four interaction frameworks: letting agents play multiple times to build relationships, tracking agents' public reputations, introducing automated mediators to make decisions for them, and using binding contracts. We tested these frameworks by simulating over 50,000 interactions among six prominent AI models across diverse scenarios. We discovered that contracts and automated mediators are exceptionally powerful, successfully recovering the vast majority of the ideal collective benefit. Remarkably, when we simulated environments where one can adapt the AI models over time to maximize success, cooperation under these interaction designs jumped to near-perfect levels. This research provides a crucial blueprint for designing safe digital spaces where autonomous AI systems can reliably collaborate for the greater good.