P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
Abstract
The transition from symbolic manipulation to science-grade reasoning represents a pivotal frontier for Large Language Models (LLMs), with physics serving as the critical test anchor for binding abstract logic to physical reality. Physics demands that a model maintain physical consistency with the laws governing the universe, fundamentally requiring multimodal perception to ground abstract logic in reality. To bridge this visual-logical gap, we introduce P1-VL, a family of open-source vision-language models engineered for advanced scientific reasoning. Our method stabilizes post-training through Curriculum Reinforcement Learning, where task difficulty is progressively expanded, and further enhances inference with Agentic Augmentation for iterative self-verification. Evaluated on HiPhO, a rigorous benchmark of 13 physics olympiads exams from 2024–2025, our flagship \texttt{P1-VL-235B-A22B} becomes the first open-source Vision-Language Model (VLM) to secure 12 gold medals and achieves the state-of-the-art performance in the open-source models. Our agent-augmented system achieves the No.2 overall rank globally, trailing only Gemini-3-Pro. Beyond physics, P1-VL demonstrates remarkable scientific reasoning capacity and generalizability, establishing significant leads over base models in STEM benchmarks.