Pixels, Patterns, but No Poetry: To See The World like Humans

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Hongcheng, Huang, Zihao, Xu, Lin, Tang, Jingyi, Li, Xinhao, Liu, Yue, Li, Haoyang, Hu, Taihang, Lin, Minhua, Yang, Xinlong, Wu, Ge, Bi, Balong, Chen, Hongyu, Zhang, Wentao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913955220291584
author Gao, Hongcheng
Huang, Zihao
Xu, Lin
Tang, Jingyi
Li, Xinhao
Liu, Yue
Li, Haoyang
Hu, Taihang
Lin, Minhua
Yang, Xinlong
Wu, Ge
Bi, Balong
Chen, Hongyu
Zhang, Wentao
author_facet Gao, Hongcheng
Huang, Zihao
Xu, Lin
Tang, Jingyi
Li, Xinhao
Liu, Yue
Li, Haoyang
Hu, Taihang
Lin, Minhua
Yang, Xinlong
Wu, Ge
Bi, Balong
Chen, Hongyu
Zhang, Wentao
contents Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a fundamental question persists: Can Multimodal Large Language Models truly perceive the world as humans do? This paper shifts focus from reasoning to perception. Rather than constructing benchmarks specifically for reasoning, we introduce the Turing Eye Test (TET), a challenging perception-oriented benchmark comprising four diagnostic tasks that evaluate MLLMs' performance on synthetic images that humans process intuitively. Our findings reveal that state-of-the-art MLLMs exhibit catastrophic failures on our perceptual tasks trivial for humans. Both in-context learning and training on language backbone-effective for previous benchmarks-fail to improve performance on our tasks, while fine-tuning the vision tower enables rapid adaptation, suggesting that our benchmark poses challenges for vision tower generalization rather than for the knowledge and reasoning capabilities of the language backbone-a key gap between current MLLMs and human perception. We release a representative subset of TET tasks in this version, and will introduce more diverse tasks and methods to enhance visual generalization in future work.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16863
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pixels, Patterns, but No Poetry: To See The World like Humans
Gao, Hongcheng
Huang, Zihao
Xu, Lin
Tang, Jingyi
Li, Xinhao
Liu, Yue
Li, Haoyang
Hu, Taihang
Lin, Minhua
Yang, Xinlong
Wu, Ge
Bi, Balong
Chen, Hongyu
Zhang, Wentao
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a fundamental question persists: Can Multimodal Large Language Models truly perceive the world as humans do? This paper shifts focus from reasoning to perception. Rather than constructing benchmarks specifically for reasoning, we introduce the Turing Eye Test (TET), a challenging perception-oriented benchmark comprising four diagnostic tasks that evaluate MLLMs' performance on synthetic images that humans process intuitively. Our findings reveal that state-of-the-art MLLMs exhibit catastrophic failures on our perceptual tasks trivial for humans. Both in-context learning and training on language backbone-effective for previous benchmarks-fail to improve performance on our tasks, while fine-tuning the vision tower enables rapid adaptation, suggesting that our benchmark poses challenges for vision tower generalization rather than for the knowledge and reasoning capabilities of the language backbone-a key gap between current MLLMs and human perception. We release a representative subset of TET tasks in this version, and will introduce more diverse tasks and methods to enhance visual generalization in future work.
title Pixels, Patterns, but No Poetry: To See The World like Humans
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.16863