QuantumSemEval: Benchmarking LLMs' Understanding of Quantum Programs via Semantic Equivalence Checking
Niloy Kumar Mondal
Abstract
Large language models (LLMs) have shown promise in reasoning about code, yet their ability to perform semantic reasoning over quantum programs remains largely unexplored. To address this gap, we introduce $\textbf{QuantumSemEval}$, a benchmark designed to evaluate LLMs’ capacity to determine semantic equivalence between pairs of quantum programs, comprising 260 labeled pairs (130 equivalent and 130 non-equivalent) across diverse algorithms. We evaluate thirteen LLMs spanning proprietary and open-weight architectures. While models reliably detect non-equivalence, identifying semantically equivalent circuits is substantially more challenging. Top performance on equivalent pairs reaches only 70.8\%, with the majority of models performing below the 50\% random baseline. These results highlight the current limitations of LLMs in semantic reasoning and provide a foundation for future work aimed at enhancing their understanding of quantum programs.
Chat is not available.
Successful Page Load