HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?
Abstract
Recently, the physics reasoning capabilities of (M)LLMs have attracted growing attention. However, existing physics benchmarks lack systematic coverage of recent physics Olympiads and direct comparison with human contestants. We present HiPhO, the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. HiPhO highlights three key innovations. (1) Comprehensive data: it compiles 13 latest Olympiads from 2024--2025, covering international and regional competitions and spanning mixed modalities from text-only to diagram-based problems. (2) Professional evaluation: it adopts official rubrics for fine-grained answer- and step-level grading aligned with human examiners. (3) Human-level comparison: it assigns gold, silver, and bronze medals to models based on official medal scores, enabling direct comparison with human contestants. Evaluating 30 (M)LLMs across 13 exams, we find that most open-source MLLMs remain at or below the bronze level, open-source LLMs demonstrate notable progress with multiple gold medals, and closed-source MLLMs achieve 6-13 gold medals, while most models still fall well short of full marks. These results underscore the substantial gap between current (M)LLMs and top human contestants, as well as the room for further improvement. The dataset and leaderboard are available at https://huggingface.co/datasets/SciYu/HiPhO and https://phyarena.github.io, respectively.
Lay Summary
Can today’s AI solve physics problems as well as the world’s best high school students? Physics Olympiad exams offer a demanding way to find out. These problems require more than plugging numbers into formulas: students must understand physical principles, interpret diagrams, and build multi-step solutions. We introduce HiPhO, a new benchmark built from 13 recent physics Olympiad exams from 2024 and 2025. HiPhO includes both text-only and diagram-based problems, and grades AI answers using the official scoring rubrics from real competitions. This means models can receive partial credit for correct reasoning steps, just as human contestants do. We also compare model scores with real medal thresholds, allowing AI models to be evaluated at gold, silver, or bronze levels. Testing 30 AI models reveals both rapid progress and clear limitations. The strongest models reach gold-medal level on many exams, showing impressive advances in scientific reasoning. Yet they still trail the top human contestants, especially on problems that require careful visual understanding or subtle physical reasoning. HiPhO provides a realistic yardstick for tracking how close AI is to human-level physics problem solving.