LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Antoine Peyronnet ⋅ Fabian Gloeckle ⋅ Amaury Hayat
Abstract
We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of contest or textbook-style problems as proxies for mathematical research. Instead, we establish an updatable benchmark evaluating models directly on the latest research results in mathematics. This consists of an automatic pipeline that extracts lemmas from the arXiv and rewrites them into self-contained statements by making all assumptions and required definitions explicit. It results in a benchmark that can be updated regularly with new problems taken directly from human mathematical research, while previous instances can be used for training without compromising future evaluations. We benchmark current state-of-the-art LLMs, which went from 10-15$\%$ accuracy to 40.8\% accuracy in theorem proving (pass@1) depending on the model, showing that there is both a very fast evolution of LLMs capabilities with respect to human research-level mathematics and still a large margin of progression for general public LLMs to reach human-level proving capabilities in a research context.
Chat is not available.
Successful Page Load