MMR: Evaluating Reading Ability of Large Multimodal Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Jian, Zhang, Ruiyi, Zhou, Yufan, Rossi, Ryan, Gu, Jiuxiang, Chen, Changyou
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913481692807168
author Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Rossi, Ryan
Gu, Jiuxiang
Chen, Changyou
author_facet Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Rossi, Ryan
Gu, Jiuxiang
Chen, Changyou
contents Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question answering, and many LMMs now easily achieve high scores. This means that current benchmarks fail to accurately reflect performance of different models, and a natural idea is to build a new benchmark to evaluate their complex reasoning and spatial understanding abilities. In this work, we propose the Multi-Modal Reading (MMR) benchmark in 11 diverse tasks to evaluate LMMs for text-rich image understanding. MMR is the first text-rich image benchmark built on human annotations with the help of language models. By evaluating several state-of-the-art LMMs, including GPT-4o, it reveals the limited capabilities of existing LMMs underscoring the value of our benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMR: Evaluating Reading Ability of Large Multimodal Models
Chen, Jian
Zhang, Ruiyi
Zhou, Yufan
Rossi, Ryan
Gu, Jiuxiang
Chen, Changyou
Computer Vision and Pattern Recognition
Large multimodal models (LMMs) have demonstrated impressive capabilities in understanding various types of image, including text-rich images. Most existing text-rich image benchmarks are simple extraction-based question answering, and many LMMs now easily achieve high scores. This means that current benchmarks fail to accurately reflect performance of different models, and a natural idea is to build a new benchmark to evaluate their complex reasoning and spatial understanding abilities. In this work, we propose the Multi-Modal Reading (MMR) benchmark in 11 diverse tasks to evaluate LMMs for text-rich image understanding. MMR is the first text-rich image benchmark built on human annotations with the help of language models. By evaluating several state-of-the-art LMMs, including GPT-4o, it reveals the limited capabilities of existing LMMs underscoring the value of our benchmark.
title MMR: Evaluating Reading Ability of Large Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.14594