SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915534426079232 |
|---|---|
| author | Xiang, Kun Li, Heng Zhang, Terry Jingchen Huang, Yinya Liu, Zirong Qu, Peixin He, Jixi Chen, Jiaqi Yuan, Yu-Jie Han, Jianhua Xu, Hang Li, Hanhui Sachan, Mrinmaya Liang, Xiaodan |
| author_facet | Xiang, Kun Li, Heng Zhang, Terry Jingchen Huang, Yinya Liu, Zirong Qu, Peixin He, Jixi Chen, Jiaqi Yuan, Yu-Jie Han, Jianhua Xu, Hang Li, Hanhui Sachan, Mrinmaya Liang, Xiaodan |
| contents | We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_19099 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning Xiang, Kun Li, Heng Zhang, Terry Jingchen Huang, Yinya Liu, Zirong Qu, Peixin He, Jixi Chen, Jiaqi Yuan, Yu-Jie Han, Jianhua Xu, Hang Li, Hanhui Sachan, Mrinmaya Liang, Xiaodan Artificial Intelligence Physics Education Popular Physics We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts. |
| title | SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning |
| topic | Artificial Intelligence Physics Education Popular Physics |
| url | https://arxiv.org/abs/2505.19099 |