One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective
Juan Duque ⋅ Sergio García Heredia ⋅ Vinicius Hernandes ⋅ Eliska Greplova ⋅ Thomas Spriggs ⋅ Aaron Courville ⋅ Anna Dawid
Abstract
Neural quantum states (NQS) provides a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov chain methods. Yet their optimization remains comparatively underexplored: Adam is a scalable method but ignores function space geometry, while stochastic reconfiguration is principled but costly and numerically fragile in large models. To address this gap, we show that variational energy minimization can be viewed as an advantage policy-gradient problem over the Born distribution, motivating trust-region optimization for NQS training. We introduce \emph{Proximal Wavefunction Optimization} (PWO), a trust-region algorithm that clips probability-ratio changes in the amplitude channel and wrapped phase increments in the phase channel. PWO avoids explicit matrix inversion, reuses samples across inner updates, and preserves the scalability of first-order optimization. Across Ising, Heisenberg, and frustrated $J_1-J_2$ spin chains, PWO improves stability and wall-clock convergence over Adam, minSR, and SPRING. Finally, we fine-tune a $1.5$B-parameter RWKV-7 model as a neural quantum state, demonstrating NQS optimization at a scale over three orders of magnitude beyond prior work.
Chat is not available.
Successful Page Load