PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915291289616384 |
|---|---|
| author | Qiu, Shi Guo, Shaoyang Song, Zhuo-Yang Sun, Yunbo Cai, Zeyu Wei, Jiashen Luo, Tianyu Yin, Yixuan Zhang, Haoxu Hu, Yi Wang, Chenyang Tang, Chencheng Chang, Haoling Liu, Qi Zhou, Ziheng Zhang, Tianyu Zhang, Jingtian Liu, Zhangyi Li, Minghao Zhang, Yuku Jing, Boxuan Yin, Xianqi Ren, Yutong Fu, Zizhuo Ji, Jiaming Wang, Weike Tian, Xudong Lv, Anqi Man, Laifu Li, Jianxiang Tao, Feiyu Sun, Qihua Liang, Zhou Mu, Yushu Li, Zhongxuan Zhang, Jing-Jun Zhang, Shutao Li, Xiaotian Xia, Xingqi Lin, Jiawei Shen, Zheyu Chen, Jiahang Xiong, Qiuhao Wang, Binran Wang, Fengyuan Ni, Ziyang Zhang, Bohan Cui, Fan Shao, Changkun Cao, Qing-Hong Luo, Ming-xing Yang, Yaodong Zhang, Muhan Zhu, Hua Xing |
| author_facet | Qiu, Shi Guo, Shaoyang Song, Zhuo-Yang Sun, Yunbo Cai, Zeyu Wei, Jiashen Luo, Tianyu Yin, Yixuan Zhang, Haoxu Hu, Yi Wang, Chenyang Tang, Chencheng Chang, Haoling Liu, Qi Zhou, Ziheng Zhang, Tianyu Zhang, Jingtian Liu, Zhangyi Li, Minghao Zhang, Yuku Jing, Boxuan Yin, Xianqi Ren, Yutong Fu, Zizhuo Ji, Jiaming Wang, Weike Tian, Xudong Lv, Anqi Man, Laifu Li, Jianxiang Tao, Feiyu Sun, Qihua Liang, Zhou Mu, Yushu Li, Zhongxuan Zhang, Jing-Jun Zhang, Shutao Li, Xiaotian Xia, Xingqi Lin, Jiawei Shen, Zheyu Chen, Jiahang Xiong, Qiuhao Wang, Binran Wang, Fengyuan Ni, Ziyang Zhang, Bohan Cui, Fan Shao, Changkun Cao, Qing-Hong Luo, Ming-xing Yang, Yaodong Zhang, Muhan Zhu, Hua Xing |
| contents | Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_16074 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models Qiu, Shi Guo, Shaoyang Song, Zhuo-Yang Sun, Yunbo Cai, Zeyu Wei, Jiashen Luo, Tianyu Yin, Yixuan Zhang, Haoxu Hu, Yi Wang, Chenyang Tang, Chencheng Chang, Haoling Liu, Qi Zhou, Ziheng Zhang, Tianyu Zhang, Jingtian Liu, Zhangyi Li, Minghao Zhang, Yuku Jing, Boxuan Yin, Xianqi Ren, Yutong Fu, Zizhuo Ji, Jiaming Wang, Weike Tian, Xudong Lv, Anqi Man, Laifu Li, Jianxiang Tao, Feiyu Sun, Qihua Liang, Zhou Mu, Yushu Li, Zhongxuan Zhang, Jing-Jun Zhang, Shutao Li, Xiaotian Xia, Xingqi Lin, Jiawei Shen, Zheyu Chen, Jiahang Xiong, Qiuhao Wang, Binran Wang, Fengyuan Ni, Ziyang Zhang, Bohan Cui, Fan Shao, Changkun Cao, Qing-Hong Luo, Ming-xing Yang, Yaodong Zhang, Muhan Zhu, Hua Xing Computation and Language Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed evaluation items. These deficiencies necessitate more rigorous assessment methods. To address these limitations, we introduce PHYBench, a benchmark of 500 original physics problems ranging from high school to Physics Olympiad difficulty. PHYBench addresses data contamination through original content and employs a systematic curation pipeline to eliminate flawed items. Evaluations show that PHYBench activates more tokens and provides stronger differentiation between reasoning models compared to other baselines like AIME 2024, OlympiadBench and GPQA. Even the best-performing model, Gemini 2.5 Pro, achieves only 36.9% accuracy compared to human experts' 61.9%. To further enhance evaluation precision, we introduce the Expression Edit Distance (EED) Score for mathematical expression assessment, which improves sample efficiency by 204% over binary scoring. Moreover, PHYBench effectively elicits multi-step and multi-condition reasoning, providing a platform for examining models' reasoning robustness, preferences, and deficiencies. The benchmark results and dataset are publicly available at https://www.phybench.cn/. |
| title | PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2504.16074 |