QuantumLean-Bench: A Unified Benchmark for Informal and Formal Quantum Reasoning
Abstract
Existing Large Language Model benchmarks typically evaluate either informal reasoning or formal proof generation, and are often limited to just one subject area or type of task. To address this gap in robust benchmarks, we introduce QuantumLean-Bench, a unified benchmark of 932 problems spanning chemistry, computing, cryptography, information theory, physics, and quantum machine learning. Each problem pairs an informal natural-language response with a formal counterpart in Lean, enabling direct evaluation of whether models preserve mathematical structure across reasoning modalities. We evaluate several frontier and open-source models, including GPT-5.4 and Gemini 3.1, using a task-aware 0/1/2 rubric and report both coverage and full-benchmark performance. We find that while frontier models achieve near-ceiling performance on informal reasoning (GPT-5.4: 0.981 normalized average, 96.1\% strict accuracy), performance drops substantially in the formal setting (0.621 normalized average, 36.1\% strict accuracy). These results highlight a persistent gap in current model capabilities: strong natural-language reasoning does not reliably translate into precise formal representations.