LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yanzhe, Zhang, Ruiyi, Gu, Jiuxiang, Zhou, Yufan, Lipka, Nedim, Yang, Diyi, Sun, Tong
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916112928604160
author Zhang, Yanzhe
Zhang, Ruiyi
Gu, Jiuxiang
Zhou, Yufan
Lipka, Nedim
Yang, Diyi
Sun, Tong
author_facet Zhang, Yanzhe
Zhang, Ruiyi
Gu, Jiuxiang
Zhou, Yufan
Lipka, Nedim
Yang, Diyi
Sun, Tong
contents Instruction tuning unlocks the superior capability of Large Language Models (LLM) to interact with humans. Furthermore, recent instruction-following datasets include images as visual inputs, collecting responses for image-based instructions. However, visual instruction-tuned models cannot comprehend textual details within images well. This work enhances the current visual instruction tuning pipeline with text-rich images (e.g., movie posters, book covers, etc.). Specifically, we first use publicly available OCR tools to collect results on 422K text-rich images from the LAION dataset. Moreover, we prompt text-only GPT-4 with recognized texts and image captions to generate 16K conversations, each containing question-answer pairs for text-rich images. By combining our collected data with previous multi-modal instruction-following data, our model, LLaVAR, substantially improves the LLaVA model's capability on text-based VQA datasets (up to 20% accuracy improvement) while achieving an accuracy of 91.42% on ScienceQA. The GPT-4-based instruction-following evaluation also demonstrates the improvement of our model on both natural images and text-rich images. Through qualitative analysis, LLaVAR shows promising interaction (e.g., reasoning, writing, and elaboration) skills with humans based on the latest real-world online content that combines text and images. We make our code/data/models publicly available at https://llavar.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2306_17107
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Zhang, Yanzhe
Zhang, Ruiyi
Gu, Jiuxiang
Zhou, Yufan
Lipka, Nedim
Yang, Diyi
Sun, Tong
Computer Vision and Pattern Recognition
Computation and Language
Instruction tuning unlocks the superior capability of Large Language Models (LLM) to interact with humans. Furthermore, recent instruction-following datasets include images as visual inputs, collecting responses for image-based instructions. However, visual instruction-tuned models cannot comprehend textual details within images well. This work enhances the current visual instruction tuning pipeline with text-rich images (e.g., movie posters, book covers, etc.). Specifically, we first use publicly available OCR tools to collect results on 422K text-rich images from the LAION dataset. Moreover, we prompt text-only GPT-4 with recognized texts and image captions to generate 16K conversations, each containing question-answer pairs for text-rich images. By combining our collected data with previous multi-modal instruction-following data, our model, LLaVAR, substantially improves the LLaVA model's capability on text-based VQA datasets (up to 20% accuracy improvement) while achieving an accuracy of 91.42% on ScienceQA. The GPT-4-based instruction-following evaluation also demonstrates the improvement of our model on both natural images and text-rich images. Through qualitative analysis, LLaVAR shows promising interaction (e.g., reasoning, writing, and elaboration) skills with humans based on the latest real-world online content that combines text and images. We make our code/data/models publicly available at https://llavar.github.io/.
title LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2306.17107