Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Yinghui, Kuang, Jiayi, Xing, Peng, Liu, Daixian, Zhang, Yongheng, Dong, Junnan, Guo, Shu-Yu, Li, Yangning, Zhou, Qingyu, Jiang, Wenhao, Zheng, Hai-Tao, Shen, Ying, Lin, Liang, Yu, Philip S.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908948871774208
author Li, Yinghui
Kuang, Jiayi
Xing, Peng
Liu, Daixian
Zhang, Yongheng
Dong, Junnan
Guo, Shu-Yu
Li, Yangning
Zhou, Qingyu
Jiang, Wenhao
Zheng, Hai-Tao
Shen, Ying
Lin, Liang
Yu, Philip S.
author_facet Li, Yinghui
Kuang, Jiayi
Xing, Peng
Liu, Daixian
Zhang, Yongheng
Dong, Junnan
Guo, Shu-Yu
Li, Yangning
Zhou, Qingyu
Jiang, Wenhao
Zheng, Hai-Tao
Shen, Ying
Lin, Liang
Yu, Philip S.
contents Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathematics, physics and chemistry, organized into three cognitive levels: perception and recognition, combination and reasoning, and association and critical thinking. Across leading MLLMs, we observe a consistent cognitive mismatch. Models frequently underperform on elementary symbol recognition while appearing relatively competent on more complex reasoning tasks. This recognition-reasoning inversion indicates that current systems often compensate with linguistic priors, template retrieval or procedural reasoning instead of robust visual grounding. The pattern is especially clear for sparse, low-redundancy symbols such as handwritten characters, formula graphs, circuit diagrams and chemical structures. These results show that symbolic understanding remains a major bottleneck for multimodal intelligence and motivate training and evaluation schemes that prioritize grounded perception in discrete semantic spaces.
format Preprint
id arxiv_https___arxiv_org_abs_2603_18472
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
Li, Yinghui
Kuang, Jiayi
Xing, Peng
Liu, Daixian
Zhang, Yongheng
Dong, Junnan
Guo, Shu-Yu
Li, Yangning
Zhou, Qingyu
Jiang, Wenhao
Zheng, Hai-Tao
Shen, Ying
Lin, Liang
Yu, Philip S.
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathematics, physics and chemistry, organized into three cognitive levels: perception and recognition, combination and reasoning, and association and critical thinking. Across leading MLLMs, we observe a consistent cognitive mismatch. Models frequently underperform on elementary symbol recognition while appearing relatively competent on more complex reasoning tasks. This recognition-reasoning inversion indicates that current systems often compensate with linguistic priors, template retrieval or procedural reasoning instead of robust visual grounding. The pattern is especially clear for sparse, low-redundancy symbols such as handwritten characters, formula graphs, circuit diagrams and chemical structures. These results show that symbolic understanding remains a major bottleneck for multimodal intelligence and motivate training and evaluation schemes that prioritize grounded perception in discrete semantic spaces.
title Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.18472