CoffeeBench: A Benchmark for Long-Horizon Strategic Decision-Making in Multi-Agent Economies
Abstract
As LLM agents advance beyond short-horizon tasks, evaluating their ability to perform long-horizon strategic decision-making in realistic multi-agent settings remains a key challenge. We introduce \textbf{CoffeeBench}, a benchmark for evaluating LLM agents in a dynamic multi-agent economic environment with farmers, roasters, and retailers forming a multi-stage supply chain. The environment models firms interacting over a multi-month horizon with evolving supply, demand, and pricing, creating sustained competition and supply-chain dynamics. Each agent operates autonomously to maximize net income through communication, negotiation, and trading with other agents. We evaluate several recent LLMs on CoffeeBench by assigning each evaluated model the role of a roaster. We find that stronger models proactively engage in trading, negotiation, and communication with counterparties, whereas weaker models behave more reactively with limited interaction. We further observe thatClaude~Haiku~4.5 exhibits an \emph{idle-drift} failure mode in which the agent maintains coherent reasoning traces but repeatedly remains inactive, resulting in prolonged operational inactivity and low net income.