WF-Bench: A Benchmark for Neural-Network WaveFunction Expressivity and Scaling Laws
Abstract
We present a comprehensive benchmarking dataset and empirical scaling law analysis for neural network wavefunctions by matching them to a wide spectrum of famous many body target wavefunctions. The dataset, WF-Bench, spans multiple distinct regimes of strongly correlated quantum matter, including topological states, Wigner crystals, and superconducting wavefunctions, providing a diverse and challenging test bed for neural network wavefunction expressivity. We introduce a systematic and reproducible benchmarking protocol for target wavefunction matching, enabling consistent performance evaluation across different neural network wavefunction architectures. By using wavefunction fidelity as the uniform metric, we discover empirical scaling laws that characterize how representability depends on system size and key model parameters, including number of determinant and model depth. By applying our benchmark protocol on Psiformer and Ferminet, we show that WF-Bench establishes a unified dataset driven framework for evaluating and comparing neural network wavefunctions and for guiding the design of future architectures.
Lay Summary
An important application of AI in quantum physics is representing many-body quantum wavefunctions. However, we still lack a clear understanding of which neural network architectures can accurately capture different types of quantum systems. This makes it difficult to compare methods fairly or to know when a neural-network wavefunction is expressive enough for a given problem. We introduce WF-Bench, a benchmark for testing how well neural networks can reproduce important quantum wavefunctions. These target wavefunctions come from topological matter, superconductors, and Wigner crystals, which represent very different physical regimes. Rather than evaluating neural networks only through energy, we directly measure how closely the learned wavefunctions match known target wavefunctions. Using WF-Bench, we compare two widely used neural-network wavefunction architectures, FermiNet and PsiFormer. We find that matching accuracy decreases with system size in a regular empirical pattern. We also find that increasing model size improves accuracy substantially at first, but gives smaller gains once the model becomes sufficiently large. Our results provide a practical way to compare neural-network wavefunctions and can help guide the design of future AI architectures for quantum physics and chemistry.