QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Abstract
Lay Summary
AI systems can now write impressive-looking math proofs, but it is still unclear whether other AI systems can reliably judge whether those proofs are correct. This paper introduces QEDBENCH, a testbed for checking AI graders on university-level math proofs. The benchmark uses 272 expert-curated proof problems, more than 1,300 AI-generated proofs, seven AI judge models, five AI solver models, and over 1,000 hours of evaluation by PhD-level experts. The study finds that many AI judges are too generous: they often give high scores to proofs that look polished but contain serious logical mistakes. One model, Llama 4 Maverick, passed 90.2% of solutions, while human experts passed only 67.7%. The hardest failures appear in areas such as combinatorics and graph theory, where success requires building a careful argument rather than following a familiar recipe. The paper also shows that stricter written rubrics do not fully solve the problem, because AI judges often keep their built-in grading habits. Overall, QEDBENCH argues that progress in AI mathematics needs better proof-checking, not just better proof-writing, especially before automated graders are trusted in education or research.