MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shan, Bin, Fei, Xiang, Shi, Wei, Wang, An-Lan, Tang, Guozhi, Liao, Lei, Tang, Jingqun, Bai, Xiang, Huang, Can
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909350296027136
author Shan, Bin
Fei, Xiang
Shi, Wei
Wang, An-Lan
Tang, Guozhi
Liao, Lei
Tang, Jingqun
Bai, Xiang
Huang, Can
author_facet Shan, Bin
Fei, Xiang
Shi, Wei
Wang, An-Lan
Tang, Guozhi
Liao, Lei
Tang, Jingqun
Bai, Xiang
Huang, Can
contents The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchmarks tailored to the scenario emphasize perceptual capabilities, while overlooking the assessment of cognitive abilities. To address this limitation, we introduce a Multimodal benchmark towards Text-rich visual scenes, to evaluate the Cognitive capabilities of MLLMs through visual reasoning and content-creation tasks (MCTBench). To mitigate potential evaluation bias from the varying distributions of datasets, MCTBench incorporates several perception tasks (e.g., scene text recognition) to ensure a consistent comparison of both the cognitive and perceptual capabilities of MLLMs. To improve the efficiency and fairness of content-creation evaluation, we conduct an automatic evaluation pipeline. Evaluations of various MLLMs on MCTBench reveal that, despite their impressive perceptual capabilities, their cognition abilities require enhancement. We hope MCTBench will offer the community an efficient resource to explore and enhance cognitive capabilities towards text-rich visual scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
Shan, Bin
Fei, Xiang
Shi, Wei
Wang, An-Lan
Tang, Guozhi
Liao, Lei
Tang, Jingqun
Bai, Xiang
Huang, Can
Computer Vision and Pattern Recognition
The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchmarks tailored to the scenario emphasize perceptual capabilities, while overlooking the assessment of cognitive abilities. To address this limitation, we introduce a Multimodal benchmark towards Text-rich visual scenes, to evaluate the Cognitive capabilities of MLLMs through visual reasoning and content-creation tasks (MCTBench). To mitigate potential evaluation bias from the varying distributions of datasets, MCTBench incorporates several perception tasks (e.g., scene text recognition) to ensure a consistent comparison of both the cognitive and perceptual capabilities of MLLMs. To improve the efficiency and fairness of content-creation evaluation, we conduct an automatic evaluation pipeline. Evaluations of various MLLMs on MCTBench reveal that, despite their impressive perceptual capabilities, their cognition abilities require enhancement. We hope MCTBench will offer the community an efficient resource to explore and enhance cognitive capabilities towards text-rich visual scenes.
title MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.11538