Toward Unified Evaluation of Mathematical AI: A Survey and Framework
Abstract
Large language models have achieved remarkable results on mathematical reasoning tasks, yet the field lacks a unified framework for evaluating both natural-language solutions and formal proofs. Existing surveys treat formal verification, LLM-as-judge, process reward models, and human grading as isolated paradigms organized by training pipeline or dataset role, while the 2024–2026 period has produced a fragmented landscape of over thirty new benchmarks spanning competition mathematics, undergraduate coursework, and research-level problems. We present a paradigm centric taxonomy crossing the formal/informal boundary, organizing evaluation methodology into six paradigms—symbolic checkers, formal proof verifiers, LLM-as-judge, process/outcome reward models, human expert grading, and hybrid generative verifiers—and systematically classify benchmarks along five orthogonal axes: difficulty tier, answer type (final-answer vs. proof), formalization level, contamination resistance, and linguistic breadth. We then provide direct evidence that paradigms diverge systematically: a synthesis of published results and a confirmatory study on the Open Proof Corpus (n=150) yield cross-paradigm Cohen’s κ in the 0.04–0.28 range, with correlated over-acceptance between LLM- judge and PRM signals in 48% of disagreements. Our critical gap analysis reveals that no existing verifier paradigm simultaneously satisfies correctness guarantees, scalability, and alignment with human mathematical judgment, motivating a research agenda toward unified evaluation.