QMP-Bench: A Research-Level Benchmark for Autonomous End-to-End Quantum Many-Body Simulations
Ken Deng ⋅ Xiangfei Wang ⋅ Guijing Duan ⋅ Chen Mo ⋅ Junkun Huang ⋅ Runqing Zhang ⋅ Ling Qian ⋅ Zhiguo Huang ⋅ Jize Han ⋅ Di Luo
Abstract
Although LLM-driven systems exhibit remarkable potential for automated scientific discovery, the community severely lacks realistic, end-to-end benchmark datasets for actual research, particularly for scientific simulations. To address this gap, we introduce QMP-Bench, a research-level dataset comprising $100$ end-to-end simulation tasks extracted directly from $21$ high-impact articles in quantum many-body physics. Unlike human-curated problems, QMP-Bench challenges agents to autonomously reproduce physical results from scratch using domain-specific numerical libraries including ITensors, NetKet, Qiskit, and ORCA. Baseline evaluations of frontier LLMs on these tasks reveal severe struggles with domain-specific programming errors and physical hallucinations, whereas introducing a framework equipped with verification mechanisms significantly improves the performance. Consequently, QMP-Bench serves as a rigorous testbed for driving the development of reliable and physics-informed AI physicists.
Chat is not available.
Successful Page Load