Last-Iterate Convergence of Regularized Gradient Methods for Stochastic Monotone Variational Inequalities
Abstract
Lay Summary
When multiple AI agents interact (for example, in games or in training generative AI), they often need to settle toward a stable equilibrium. A core challenge is making the current strategy good at every moment, not just on average, since each player must act now. This becomes much harder when the feedback signals are noisy, as is typical in real systems. We study two simple algorithms that update once per round from noisy feedback, and prove they steer the current strategy toward equilibrium. Crucially, our guarantees hold at every step, without knowing in advance how long the process will run. We achieve this through a careful balance of a shrinking pull toward a reference point and a decreasing step size. Earlier work either required knowing the run time up front or only guaranteed averages rather than the current state. Our results show that simple, natural algorithms suffice for the harder problem of staying close to equilibrium at any time. The guarantees even sharpen when the noise level is small. This may inform learning rules for AI systems with multiple interacting agents, such as in strategic games or setups where models compete and improve each other.