Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Abstract
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for tracking progress: they are often narrow in scope, quickly saturated, and rarely updated. We instead need evaluation platforms: continuously maintained systems that run, aggregate, and analyze evaluations across many benchmarks. In this work, we build on the original MathArena benchmark by broadening it from final-answer olympiad problems into an evaluation platform for mathematical reasoning with LLMs. MathArena now covers proof-based competitions, research-level arXiv problems, and formal proof generation in Lean. We also maintain a clear evaluation protocol and regularly add benchmarks as model capabilities improve. Notably, GPT-5.5 reaches 98% on the 2026 USA Math Olympiad and 74% on research-level questions, showing that frontier models can now solve extremely challenging mathematical problems. This highlights the importance of continuously maintained platforms like MathArena for tracking progress in mathematical reasoning.