Exploring the Capabilities of Large Multimodal Models on Dense Text

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shuo, Yang, Biao, Li, Zhang, Ma, Zhiyin, Liu, Yuliang, Bai, Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909197887602688
author Zhang, Shuo
Yang, Biao
Li, Zhang
Ma, Zhiyin
Liu, Yuliang
Bai, Xiang
author_facet Zhang, Shuo
Yang, Biao
Li, Zhang
Ma, Zhiyin
Liu, Yuliang
Bai, Xiang
contents While large multi-modal models (LMM) have shown notable progress in multi-modal tasks, their capabilities in tasks involving dense textual content remains to be fully explored. Dense text, which carries important information, is often found in documents, tables, and product descriptions. Understanding dense text enables us to obtain more accurate information, assisting in making better decisions. To further explore the capabilities of LMM in complex text tasks, we propose the DT-VQA dataset, with 170k question-answer pairs. In this paper, we conduct a comprehensive evaluation of GPT4V, Gemini, and various open-source LMMs on our dataset, revealing their strengths and weaknesses. Furthermore, we evaluate the effectiveness of two strategies for LMM: prompt engineering and downstream fine-tuning. We find that even with automatically labeled training datasets, significant improvements in model performance can be achieved. We hope that this research will promote the study of LMM in dense text tasks. Code will be released at https://github.com/Yuliang-Liu/MultimodalOCR.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring the Capabilities of Large Multimodal Models on Dense Text
Zhang, Shuo
Yang, Biao
Li, Zhang
Ma, Zhiyin
Liu, Yuliang
Bai, Xiang
Computation and Language
Artificial Intelligence
While large multi-modal models (LMM) have shown notable progress in multi-modal tasks, their capabilities in tasks involving dense textual content remains to be fully explored. Dense text, which carries important information, is often found in documents, tables, and product descriptions. Understanding dense text enables us to obtain more accurate information, assisting in making better decisions. To further explore the capabilities of LMM in complex text tasks, we propose the DT-VQA dataset, with 170k question-answer pairs. In this paper, we conduct a comprehensive evaluation of GPT4V, Gemini, and various open-source LMMs on our dataset, revealing their strengths and weaknesses. Furthermore, we evaluate the effectiveness of two strategies for LMM: prompt engineering and downstream fine-tuning. We find that even with automatically labeled training datasets, significant improvements in model performance can be achieved. We hope that this research will promote the study of LMM in dense text tasks. Code will be released at https://github.com/Yuliang-Liu/MultimodalOCR.
title Exploring the Capabilities of Large Multimodal Models on Dense Text
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.06706