Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Yun, Wang, Futing, Cheng, Qianjia, Yu, Fangchen, Lei, Haodi, Yan, Jianhao, Li, Chenxi, Chen, Jiacheng, Zhao, Yufeng, Wan, Haiyuan, Zhang, Yuchen, Zheng, Shenghe, Yao, Junchi, Zhang, Qingyang, He, Haonan, Zeng, Wenxuan, Sheng, Li, Xie, Chengxing, Zuo, Yuxin, Li, Yizhuo, Wu, Yulun, Huang, Rui, Zhou, Dongzhan, Chen, Kai, Qiao, Yu, Bai, Lei, Cheng, Yu, Ding, Ning, Zhou, Bowen, Ye, Peng, Cui, Ganqu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:https://arxiv.org/abs/2602.09443
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914318730133504
author Luo, Yun
Wang, Futing
Cheng, Qianjia
Yu, Fangchen
Lei, Haodi
Yan, Jianhao
Li, Chenxi
Chen, Jiacheng
Zhao, Yufeng
Wan, Haiyuan
Zhang, Yuchen
Zheng, Shenghe
Yao, Junchi
Zhang, Qingyang
He, Haonan
Zeng, Wenxuan
Sheng, Li
Xie, Chengxing
Zuo, Yuxin
Li, Yizhuo
Wu, Yulun
Huang, Rui
Zhou, Dongzhan
Chen, Kai
Qiao, Yu
Bai, Lei
Cheng, Yu
Ding, Ning
Zhou, Bowen
Ye, Peng
Cui, Ganqu
author_facet Luo, Yun
Wang, Futing
Cheng, Qianjia
Yu, Fangchen
Lei, Haodi
Yan, Jianhao
Li, Chenxi
Chen, Jiacheng
Zhao, Yufeng
Wan, Haiyuan
Zhang, Yuchen
Zheng, Shenghe
Yao, Junchi
Zhang, Qingyang
He, Haonan
Zeng, Wenxuan
Sheng, Li
Xie, Chengxing
Zuo, Yuxin
Li, Yizhuo
Wu, Yulun
Huang, Rui
Zhou, Dongzhan
Chen, Kai
Qiao, Yu
Bai, Lei
Cheng, Yu
Ding, Ning
Zhou, Bowen
Ye, Peng
Cui, Ganqu
contents The transition from symbolic manipulation to science-grade reasoning represents a pivotal frontier for Large Language Models (LLMs), with physics serving as the critical test anchor for binding abstract logic to physical reality. Physics demands that a model maintain physical consistency with the laws governing the universe, a task that fundamentally requires multimodal perception to ground abstract logic in reality. At the Olympiad level, diagrams are often constitutive rather than illustrative, containing essential constraints, such as boundary conditions and spatial symmetries, that are absent from the text. To bridge this visual-logical gap, we introduce P1-VL, a family of open-source vision-language models engineered for advanced scientific reasoning. Our method harmonizes Curriculum Reinforcement Learning, which employs progressive difficulty expansion to stabilize post-training, with Agentic Augmentation, enabling iterative self-verification at inference. Evaluated on HiPhO, a rigorous benchmark of 13 exams from 2024-2025, our flagship P1-VL-235B-A22B becomes the first open-source Vision-Language Model (VLM) to secure 12 gold medals and achieves the state-of-the-art performance in the open-source models. Our agent-augmented system achieves the No.2 overall rank globally, trailing only Gemini-3-Pro. Beyond physics, P1-VL demonstrates remarkable scientific reasoning capacity and generalizability, establishing significant leads over base models in STEM benchmarks. By open-sourcing P1-VL, we provide a foundational step toward general-purpose physical intelligence to better align visual perceptions with abstract physical laws for machine scientific discovery.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09443
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
Luo, Yun
Wang, Futing
Cheng, Qianjia
Yu, Fangchen
Lei, Haodi
Yan, Jianhao
Li, Chenxi
Chen, Jiacheng
Zhao, Yufeng
Wan, Haiyuan
Zhang, Yuchen
Zheng, Shenghe
Yao, Junchi
Zhang, Qingyang
He, Haonan
Zeng, Wenxuan
Sheng, Li
Xie, Chengxing
Zuo, Yuxin
Li, Yizhuo
Wu, Yulun
Huang, Rui
Zhou, Dongzhan
Chen, Kai
Qiao, Yu
Bai, Lei
Cheng, Yu
Ding, Ning
Zhou, Bowen
Ye, Peng
Cui, Ganqu
Artificial Intelligence
The transition from symbolic manipulation to science-grade reasoning represents a pivotal frontier for Large Language Models (LLMs), with physics serving as the critical test anchor for binding abstract logic to physical reality. Physics demands that a model maintain physical consistency with the laws governing the universe, a task that fundamentally requires multimodal perception to ground abstract logic in reality. At the Olympiad level, diagrams are often constitutive rather than illustrative, containing essential constraints, such as boundary conditions and spatial symmetries, that are absent from the text. To bridge this visual-logical gap, we introduce P1-VL, a family of open-source vision-language models engineered for advanced scientific reasoning. Our method harmonizes Curriculum Reinforcement Learning, which employs progressive difficulty expansion to stabilize post-training, with Agentic Augmentation, enabling iterative self-verification at inference. Evaluated on HiPhO, a rigorous benchmark of 13 exams from 2024-2025, our flagship P1-VL-235B-A22B becomes the first open-source Vision-Language Model (VLM) to secure 12 gold medals and achieves the state-of-the-art performance in the open-source models. Our agent-augmented system achieves the No.2 overall rank globally, trailing only Gemini-3-Pro. Beyond physics, P1-VL demonstrates remarkable scientific reasoning capacity and generalizability, establishing significant leads over base models in STEM benchmarks. By open-sourcing P1-VL, we provide a foundational step toward general-purpose physical intelligence to better align visual perceptions with abstract physical laws for machine scientific discovery.
title P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
topic Artificial Intelligence
url https://arxiv.org/abs/2602.09443