PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lintao, Su, Encheng, Liu, Jiaqi, Li, Pengze, Xiao, Jiabei, Zhang, Wenlong, Dai, Xinnan, Chen, Xi, Meng, Yuan, Bai, Lei, Ouyang, Wanli, Tang, Shixiang, Wang, Aoran, Ma, Xinzhu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915789477511168
author Wang, Lintao
Su, Encheng
Liu, Jiaqi
Li, Pengze
Xiao, Jiabei
Zhang, Wenlong
Dai, Xinnan
Chen, Xi
Meng, Yuan
Bai, Lei
Ouyang, Wanli
Tang, Shixiang
Wang, Aoran
Ma, Xinzhu
author_facet Wang, Lintao
Su, Encheng
Liu, Jiaqi
Li, Pengze
Xiao, Jiabei
Zhang, Wenlong
Dai, Xinnan
Chen, Xi
Meng, Yuan
Bai, Lei
Ouyang, Wanli
Tang, Shixiang
Wang, Aoran
Ma, Xinzhu
contents Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogical assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagram. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current models encounter substantial challenges in physics reasoning, where GPT-5 achieves only 51.6% accuracy in the PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
Wang, Lintao
Su, Encheng
Liu, Jiaqi
Li, Pengze
Xiao, Jiabei
Zhang, Wenlong
Dai, Xinnan
Chen, Xi
Meng, Yuan
Bai, Lei
Ouyang, Wanli
Tang, Shixiang
Wang, Aoran
Ma, Xinzhu
Artificial Intelligence
Physics problem-solving is a challenging domain for AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogical assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagram. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current models encounter substantial challenges in physics reasoning, where GPT-5 achieves only 51.6% accuracy in the PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding.
title PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
topic Artificial Intelligence
url https://arxiv.org/abs/2506.17667