Statistical Complexity of Soft Bellman Residual Minimization
Kyoungseok Jang
Abstract
Bellman residual minimization (BRM) provides a scalable, gradient-based approach to offline reinforcement learning in large state spaces. While globally convergent gradient-based soft (i.e., entropy-regularized) BRM methods for neural networks have recently been established, their statistical complexity under stochastic gradient descent remains largely unknown. In this paper, we address this theoretical gap for soft BRM. Through a novel Lyapunov-based analysis, we establish an $\mathcal{O}(1 / n)$ average argument stability bound, which translates directly into a $\mathcal{O}(1 / n)$ statistical complexity for the soft BRM objective.
Chat is not available.
Successful Page Load