PhyX: Does Your Model Have the "Wits" for Physical Reasoning?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Hui, Wu, Taiqiang, Han, Qi, Hsieh, Yunta, Wang, Jizhou, Zhang, Yuyue, Cheng, Yuxin, Hao, Zijian, Ni, Yuansheng, Wang, Xin, Wan, Zhongwei, Zhang, Kai, Xu, Wendong, Xiong, Jing, Luo, Ping, Chen, Wenhu, Tao, Chaofan, Mao, Zhuoqing, Wong, Ngai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913865785147392
author Shen, Hui
Wu, Taiqiang
Han, Qi
Hsieh, Yunta
Wang, Jizhou
Zhang, Yuyue
Cheng, Yuxin
Hao, Zijian
Ni, Yuansheng
Wang, Xin
Wan, Zhongwei
Zhang, Kai
Xu, Wendong
Xiong, Jing
Luo, Ping
Chen, Wenhu
Tao, Chaofan
Mao, Zhuoqing
Wong, Ngai
author_facet Shen, Hui
Wu, Taiqiang
Han, Qi
Hsieh, Yunta
Wang, Jizhou
Zhang, Yuyue
Cheng, Yuxin
Hao, Zijian
Ni, Yuansheng
Wang, Xin
Wan, Zhongwei
Zhang, Kai
Xu, Wendong
Xiong, Jing
Luo, Ping
Chen, Wenhu
Tao, Chaofan
Mao, Zhuoqing
Wong, Ngai
contents Existing benchmarks fail to capture a crucial aspect of intelligence: physical reasoning, the integrated ability to combine domain knowledge, symbolic reasoning, and understanding of real-world constraints. To address this gap, we introduce PhyX: the first large-scale benchmark designed to assess models capacity for physics-grounded reasoning in visual scenarios. PhyX includes 3K meticulously curated multimodal questions spanning 6 reasoning types across 25 sub-domains and 6 core physics domains: thermodynamics, electromagnetism, mechanics, modern physics, optics, and wave\&acoustics. In our comprehensive evaluation, even state-of-the-art models struggle significantly with physical reasoning. GPT-4o, Claude3.7-Sonnet, and GPT-o4-mini achieve only 32.5%, 42.2%, and 45.8% accuracy respectively-performance gaps exceeding 29% compared to human experts. Our analysis exposes critical limitations in current models: over-reliance on memorized disciplinary knowledge, excessive dependence on mathematical formulations, and surface-level visual pattern matching rather than genuine physical understanding. We provide in-depth analysis through fine-grained statistics, detailed case studies, and multiple evaluation paradigms to thoroughly examine physical reasoning capabilities. To ensure reproducibility, we implement a compatible evaluation protocol based on widely-used toolkits such as VLMEvalKit, enabling one-click evaluation. More details are available on our project page: https://phyx-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15929
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Shen, Hui
Wu, Taiqiang
Han, Qi
Hsieh, Yunta
Wang, Jizhou
Zhang, Yuyue
Cheng, Yuxin
Hao, Zijian
Ni, Yuansheng
Wang, Xin
Wan, Zhongwei
Zhang, Kai
Xu, Wendong
Xiong, Jing
Luo, Ping
Chen, Wenhu
Tao, Chaofan
Mao, Zhuoqing
Wong, Ngai
Artificial Intelligence
Existing benchmarks fail to capture a crucial aspect of intelligence: physical reasoning, the integrated ability to combine domain knowledge, symbolic reasoning, and understanding of real-world constraints. To address this gap, we introduce PhyX: the first large-scale benchmark designed to assess models capacity for physics-grounded reasoning in visual scenarios. PhyX includes 3K meticulously curated multimodal questions spanning 6 reasoning types across 25 sub-domains and 6 core physics domains: thermodynamics, electromagnetism, mechanics, modern physics, optics, and wave\&acoustics. In our comprehensive evaluation, even state-of-the-art models struggle significantly with physical reasoning. GPT-4o, Claude3.7-Sonnet, and GPT-o4-mini achieve only 32.5%, 42.2%, and 45.8% accuracy respectively-performance gaps exceeding 29% compared to human experts. Our analysis exposes critical limitations in current models: over-reliance on memorized disciplinary knowledge, excessive dependence on mathematical formulations, and surface-level visual pattern matching rather than genuine physical understanding. We provide in-depth analysis through fine-grained statistics, detailed case studies, and multiple evaluation paradigms to thoroughly examine physical reasoning capabilities. To ensure reproducibility, we implement a compatible evaluation protocol based on widely-used toolkits such as VLMEvalKit, enabling one-click evaluation. More details are available on our project page: https://phyx-bench.github.io/.
title PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
topic Artificial Intelligence
url https://arxiv.org/abs/2505.15929