Agentic AI for Mathematical Research
Abstract
Can frontier AI systems make autonomous progress on research-level mathematical problems? Can mathematicians use them to accelerate their own research in substantive ways? As model capabilities improve, the answer to both questions is an increasingly confident, though still cautious, “yes.” In this talk, I will survey the mathematical reasoning capabilities of current frontier models, with a particular focus on agentic systems for theorem proving and mathematical exploration. I will discuss what these systems can already do, where they still fail, and what kinds of workflows seem most promising for human–AI collaboration. I will then examine the results from the First Proof, Second Batch competition in detail, with a focus on the design and performance of ProofCouncil, an agentic system for research-level mathematics.