Deep Reinforcement Learning Finds Bayes-Nash Equilibrium in Competitive Newsvendor Problems
Abstract
We investigate learning dynamics in competitive newsvendor games, a class of continuous-action games with strategic substitutes. Despite established equilibrium properties, convergence of independent learning algorithms in repeated general-sum play remains uncertain. We analyze structural properties under complete and incomplete information, deriving closed-form equilibria for a symmetric complete-information benchmark with perfect substitution. Our main theoretical contribution proves strict monotonicity in both complete-information and Bayesian models with private costs, ensuring equilibrium uniqueness and ruling out unstable dynamics. This provides convergence guarantees for variational-inequality-based algorithms. Numerical experiments using deep reinforcement learning agents with Proximal Policy Optimization empirically demonstrate convergence to Nash and Bayesian Nash equilibria, verified by equilibrium checks. These results establish a foundation for applying deep reinforcement learning in competitive inventory management.
Lay Summary
Modern companies increasingly use AI systems to make decisions in competitive environments such as pricing, advertising, and inventory management. However, when multiple learning agents adapt at the same time, learning can become unstable and fail to converge to reliable strategies. We study this challenge using a competitive inventory management problem in which firms decide how much stock to order under uncertain demand while competing for customers. Our main theoretical result shows that this interaction has a special mathematical structure that guarantees a unique desirable outcome and that it will be found by learning agents reliably. On the empirical side, we do two things. First, we verify our theoretical results in settings that possess the special structural properties. Second, we show that deep reinforcement learning agents reliably find desirable outcomes even in settings that go beyond the currently established theoretical guarantees. These findings help explain why reinforcement learning succeeds in competitive inventory problems, and they provide a foundation for more reliable AI systems in supply chains and other multi-agent decision settings.