_version_ 1866915291289616384
author Qiu, Shi
Guo, Shaoyang
Song, Zhuo-Yang
Sun, Yunbo
Cai, Zeyu
Wei, Jiashen
Luo, Tianyu
Yin, Yixuan
Zhang, Haoxu
Hu, Yi
Wang, Chenyang
Tang, Chencheng
Chang, Haoling
Liu, Qi
Zhou, Ziheng
Zhang, Tianyu
Zhang, Jingtian
Liu, Zhangyi
Li, Minghao
Zhang, Yuku
Jing, Boxuan
Yin, Xianqi
Ren, Yutong
Fu, Zizhuo
Ji, Jiaming
Wang, Weike
Tian, Xudong
Lv, Anqi
Man, Laifu
Li, Jianxiang
Tao, Feiyu
Sun, Qihua
Liang, Zhou
Mu, Yushu
Li, Zhongxuan
Zhang, Jing-Jun
Zhang, Shutao
Li, Xiaotian
Xia, Xingqi
Lin, Jiawei
Shen, Zheyu
Chen, Jiahang
Xiong, Qiuhao
Wang, Binran
Wang, Fengyuan
Ni, Ziyang
Zhang, Bohan
Cui, Fan
Shao, Changkun
Cao, Qing-Hong
Luo, Ming-xing
Yang, Yaodong
Zhang, Muhan
Zhu, Hua Xing
author_facet Qiu, Shi
Guo, Shaoyang
Song, Zhuo-Yang
Sun, Yunbo
Cai, Zeyu
Wei, Jiashen
Luo, Tianyu
Yin, Yixuan
Zhang, Haoxu
Hu, Yi
Wang, Chenyang
Tang, Chencheng
Chang, Haoling
Liu, Qi
Zhou, Ziheng
Zhang, Tianyu
Zhang, Jingtian
Liu, Zhangyi
Li, Minghao
Zhang, Yuku
Jing, Boxuan
Yin, Xianqi
Ren, Yutong
Fu, Zizhuo
Ji, Jiaming
Wang, Weike
Tian, Xudong
Lv, Anqi
Man, Laifu
Li, Jianxiang
Tao, Feiyu
Sun, Qihua
Liang, Zhou
Mu, Yushu
Li, Zhongxuan
Zhang, Jing-Jun
Zhang, Shutao
Li, Xiaotian
Xia, Xingqi
Lin, Jiawei
Shen, Zheyu
Chen, Jiahang
Xiong, Qiuhao
Wang, Binran
Wang, Fengyuan
Ni, Ziyang
Zhang, Bohan
Cui, Fan
Shao, Changkun
Cao, Qing-Hong
Luo, Ming-xing
Yang, Yaodong
Zhang, Muhan
Zhu, Hua Xing
contents Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Qiu, Shi
Guo, Shaoyang
Song, Zhuo-Yang
Sun, Yunbo
Cai, Zeyu
Wei, Jiashen
Luo, Tianyu
Yin, Yixuan
Zhang, Haoxu
Hu, Yi
Wang, Chenyang
Tang, Chencheng
Chang, Haoling
Liu, Qi
Zhou, Ziheng
Zhang, Tianyu
Zhang, Jingtian
Liu, Zhangyi
Li, Minghao
Zhang, Yuku
Jing, Boxuan
Yin, Xianqi
Ren, Yutong
Fu, Zizhuo
Ji, Jiaming
Wang, Weike
Tian, Xudong
Lv, Anqi
Man, Laifu
Li, Jianxiang
Tao, Feiyu
Sun, Qihua
Liang, Zhou
Mu, Yushu
Li, Zhongxuan
Zhang, Jing-Jun
Zhang, Shutao
Li, Xiaotian
Xia, Xingqi
Lin, Jiawei
Shen, Zheyu
Chen, Jiahang
Xiong, Qiuhao
Wang, Binran
Wang, Fengyuan
Ni, Ziyang
Zhang, Bohan
Cui, Fan
Shao, Changkun
Cao, Qing-Hong
Luo, Ming-xing
Yang, Yaodong
Zhang, Muhan
Zhu, Hua Xing
Computation and Language
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/.
title PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2504.16074