SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Kun, Li, Heng, Zhang, Terry Jingchen, Huang, Yinya, Liu, Zirong, Qu, Peixin, He, Jixi, Chen, Jiaqi, Yuan, Yu-Jie, Han, Jianhua, Xu, Hang, Li, Hanhui, Sachan, Mrinmaya, Liang, Xiaodan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915534426079232
author Xiang, Kun
Li, Heng
Zhang, Terry Jingchen
Huang, Yinya
Liu, Zirong
Qu, Peixin
He, Jixi
Chen, Jiaqi
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Li, Hanhui
Sachan, Mrinmaya
Liang, Xiaodan
author_facet Xiang, Kun
Li, Heng
Zhang, Terry Jingchen
Huang, Yinya
Liu, Zirong
Qu, Peixin
He, Jixi
Chen, Jiaqi
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Li, Hanhui
Sachan, Mrinmaya
Liang, Xiaodan
contents We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19099
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
Xiang, Kun
Li, Heng
Zhang, Terry Jingchen
Huang, Yinya
Liu, Zirong
Qu, Peixin
He, Jixi
Chen, Jiaqi
Yuan, Yu-Jie
Han, Jianhua
Xu, Hang
Li, Hanhui
Sachan, Mrinmaya
Liang, Xiaodan
Artificial Intelligence
Physics Education
Popular Physics
We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts.
title SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
topic Artificial Intelligence
Physics Education
Popular Physics
url https://arxiv.org/abs/2505.19099