Can Standard MARL Metrics Distinguish Communicative from Strategic Action?
Abstract
Multi-agent Reinforcement Learning (MARL) systems are routinely evaluated using aggregate utility metrics. A population that converges to high reward is often described as having reached "consensus". Drawing on @habermas1984, we distinguish a justified consensus from strategic action. Standard MARL objectives collapse the distinction: both record as similarly successful. We demonstrate the gap in a minimal foraging environment with one "Tyrant" agent that can unilaterally penalize peers. The system converges to high-reward equilibria nearly indistinguishable from a symmetric baseline by standard metrics. A coercion index, tracking how peers yield to credible threats, exposes what aggregate return hides. Trustworthy MARL evaluation requires diagnostics explicitly sensitive to capability asymmetry.