Advancing and Evaluating AI Agent for Frontier Physics Research
Abstract
Advances of LLMs in reasoning, domain knowledge, and tool use reveal the potential of agentic science, where AI conducts autonomous research and enables AI-driven scientific discovery. Yet such paradigm remains difficult to realize in theoretical and computational physics research, where research workflows require domain knowledge, long-horizon reasoning, and reliable numerical computation. Towards agentic physics research, we present PhysMaster, a research agent system integrating adaptive multi-trajectory exploration and trusted external knowledge to conduct reliable ultra-long-horizon exploration. Further, we construct PRL-Bench, a research-oriented benchmark adapted from 100 Physical Review Letters papers. Evaluation shows that even frontier models face frequent conceptual errors, lack of advanced theoretical-physics knowledge, and unstable derivations over extended horizons, while PhysMaster substantially mitigates these limitations. Together, PhysMaster and PRL-Bench provide a pioneering paradigm for advancing and evaluating autonomous AI scientists for frontier physics research.