TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Jingqun, Lin, Chunhui, Zhao, Zhen, Wei, Shu, Wu, Binghong, Liu, Qi, He, Yangfan, Lu, Kuan, Feng, Hao, Li, Yang, Wang, Siqi, Liao, Lei, Shi, Wei, Liu, Yuliang, Liu, Hao, Xie, Yuan, Bai, Xiang, Huang, Can
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912422863831040
author Tang, Jingqun
Lin, Chunhui
Zhao, Zhen
Wei, Shu
Wu, Binghong
Liu, Qi
He, Yangfan
Lu, Kuan
Feng, Hao
Li, Yang
Wang, Siqi
Liao, Lei
Shi, Wei
Liu, Yuliang
Liu, Hao
Xie, Yuan
Bai, Xiang
Huang, Can
author_facet Tang, Jingqun
Lin, Chunhui
Zhao, Zhen
Wei, Shu
Wu, Binghong
Liu, Qi
He, Yangfan
Lu, Kuan
Feng, Hao
Li, Yang
Wang, Siqi
Liao, Lei
Shi, Wei
Liu, Yuliang
Liu, Hao
Xie, Yuan
Bai, Xiang
Huang, Can
contents Text-centric visual question answering (VQA) has made great strides with the development of Multimodal Large Language Models (MLLMs), yet open-source models still fall short of leading models like GPT4V and Gemini, partly due to a lack of extensive, high-quality instruction tuning data. To this end, we introduce a new approach for creating a massive, high-quality instruction-tuning dataset, Square-10M, which is generated using closed-source MLLMs. The data construction process, termed Square, consists of four steps: Self-Questioning, Answering, Reasoning, and Evaluation. Our experiments with Square-10M led to three key findings: 1) Our model, TextSquare, considerably surpasses open-source previous state-of-the-art Text-centric MLLMs and sets a new standard on OCRBench(62.2%). It even outperforms top-tier models like GPT4V and Gemini in 6 of 10 text-centric benchmarks. 2) Additionally, we demonstrate the critical role of VQA reasoning data in offering comprehensive contextual insights for specific questions. This not only improves accuracy but also significantly mitigates hallucinations. Specifically, TextSquare scores an average of 75.1% across four general VQA and hallucination evaluation datasets, outperforming previous state-of-the-art models. 3) Notably, the phenomenon observed in scaling text-centric VQA datasets reveals a vivid pattern: the exponential increase of instruction tuning data volume is directly proportional to the improvement in model performance, thereby validating the necessity of the dataset scale and the high quality of Square-10M.
format Preprint
id arxiv_https___arxiv_org_abs_2404_12803
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Tang, Jingqun
Lin, Chunhui
Zhao, Zhen
Wei, Shu
Wu, Binghong
Liu, Qi
He, Yangfan
Lu, Kuan
Feng, Hao
Li, Yang
Wang, Siqi
Liao, Lei
Shi, Wei
Liu, Yuliang
Liu, Hao
Xie, Yuan
Bai, Xiang
Huang, Can
Computer Vision and Pattern Recognition
Machine Learning
Text-centric visual question answering (VQA) has made great strides with the development of Multimodal Large Language Models (MLLMs), yet open-source models still fall short of leading models like GPT4V and Gemini, partly due to a lack of extensive, high-quality instruction tuning data. To this end, we introduce a new approach for creating a massive, high-quality instruction-tuning dataset, Square-10M, which is generated using closed-source MLLMs. The data construction process, termed Square, consists of four steps: Self-Questioning, Answering, Reasoning, and Evaluation. Our experiments with Square-10M led to three key findings: 1) Our model, TextSquare, considerably surpasses open-source previous state-of-the-art Text-centric MLLMs and sets a new standard on OCRBench(62.2%). It even outperforms top-tier models like GPT4V and Gemini in 6 of 10 text-centric benchmarks. 2) Additionally, we demonstrate the critical role of VQA reasoning data in offering comprehensive contextual insights for specific questions. This not only improves accuracy but also significantly mitigates hallucinations. Specifically, TextSquare scores an average of 75.1% across four general VQA and hallucination evaluation datasets, outperforming previous state-of-the-art models. 3) Notably, the phenomenon observed in scaling text-centric VQA datasets reveals a vivid pattern: the exponential increase of instruction tuning data volume is directly proportional to the improvement in model performance, thereby validating the necessity of the dataset scale and the high quality of Square-10M.
title TextSquare: Scaling up Text-Centric Visual Instruction Tuning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2404.12803