Multi-Agent Reinforcement Learning of Karma Bidding Strategies
Abstract
Capacity-constrained shared infrastructure systems require demand management mechanisms that balance efficiency and fairness. Karma mechanisms address this challenge using an artificial, non-tradable currency that enables decentralized allocation through repeated bidding, but computing equilibrium strategies in such settings is difficult due to unknown population dynamics, stochastic demand, and computation time in practice. Here, we study suitable regimes to prove convergence towards stochastic Nash equilibria when following a multi-agent reinforcement learning approach. Computational case studies demonstrate empirically that learned policies closely approximate equilibrium behavior, and further assess impact of learning algorithm and policy initialization on convergence speed. The work highlights the practical potential of the approach for real-world decision support in repeated access-allocation settings.