Pixels, Patterns, but No Poetry: To See The World like Humans
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913955220291584 |
|---|---|
| author | Gao, Hongcheng Huang, Zihao Xu, Lin Tang, Jingyi Li, Xinhao Liu, Yue Li, Haoyang Hu, Taihang Lin, Minhua Yang, Xinlong Wu, Ge Bi, Balong Chen, Hongyu Zhang, Wentao |
| author_facet | Gao, Hongcheng Huang, Zihao Xu, Lin Tang, Jingyi Li, Xinhao Liu, Yue Li, Haoyang Hu, Taihang Lin, Minhua Yang, Xinlong Wu, Ge Bi, Balong Chen, Hongyu Zhang, Wentao |
| contents | Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a fundamental question persists: Can Multimodal Large Language Models truly perceive the world as humans do? This paper shifts focus from reasoning to perception. Rather than constructing benchmarks specifically for reasoning, we introduce the Turing Eye Test (TET), a challenging perception-oriented benchmark comprising four diagnostic tasks that evaluate MLLMs' performance on synthetic images that humans process intuitively. Our findings reveal that state-of-the-art MLLMs exhibit catastrophic failures on our perceptual tasks trivial for humans. Both in-context learning and training on language backbone-effective for previous benchmarks-fail to improve performance on our tasks, while fine-tuning the vision tower enables rapid adaptation, suggesting that our benchmark poses challenges for vision tower generalization rather than for the knowledge and reasoning capabilities of the language backbone-a key gap between current MLLMs and human perception. We release a representative subset of TET tasks in this version, and will introduce more diverse tasks and methods to enhance visual generalization in future work. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_16863 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Pixels, Patterns, but No Poetry: To See The World like Humans Gao, Hongcheng Huang, Zihao Xu, Lin Tang, Jingyi Li, Xinhao Liu, Yue Li, Haoyang Hu, Taihang Lin, Minhua Yang, Xinlong Wu, Ge Bi, Balong Chen, Hongyu Zhang, Wentao Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a fundamental question persists: Can Multimodal Large Language Models truly perceive the world as humans do? This paper shifts focus from reasoning to perception. Rather than constructing benchmarks specifically for reasoning, we introduce the Turing Eye Test (TET), a challenging perception-oriented benchmark comprising four diagnostic tasks that evaluate MLLMs' performance on synthetic images that humans process intuitively. Our findings reveal that state-of-the-art MLLMs exhibit catastrophic failures on our perceptual tasks trivial for humans. Both in-context learning and training on language backbone-effective for previous benchmarks-fail to improve performance on our tasks, while fine-tuning the vision tower enables rapid adaptation, suggesting that our benchmark poses challenges for vision tower generalization rather than for the knowledge and reasoning capabilities of the language backbone-a key gap between current MLLMs and human perception. We release a representative subset of TET tasks in this version, and will introduce more diverse tasks and methods to enhance visual generalization in future work. |
| title | Pixels, Patterns, but No Poetry: To See The World like Humans |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2507.16863 |