VIGC: Visual Instruction Generation and Correction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Bin, Wu, Fan, Han, Xiao, Peng, Jiahui, Zhong, Huaping, Zhang, Pan, Dong, Xiaoyi, Li, Weijia, Li, Wei, Wang, Jiaqi, He, Conghui
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929233727586304
author Wang, Bin
Wu, Fan
Han, Xiao
Peng, Jiahui
Zhong, Huaping
Zhang, Pan
Dong, Xiaoyi
Li, Weijia
Li, Wei
Wang, Jiaqi
He, Conghui
author_facet Wang, Bin
Wu, Fan
Han, Xiao
Peng, Jiahui
Zhong, Huaping
Zhang, Pan
Dong, Xiaoyi
Li, Weijia
Li, Wei
Wang, Jiaqi
He, Conghui
contents The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models (MLLMs) to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC.
format Preprint
id arxiv_https___arxiv_org_abs_2308_12714
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VIGC: Visual Instruction Generation and Correction
Wang, Bin
Wu, Fan
Han, Xiao
Peng, Jiahui
Zhong, Huaping
Zhang, Pan
Dong, Xiaoyi
Li, Weijia
Li, Wei
Wang, Jiaqi
He, Conghui
Computer Vision and Pattern Recognition
Artificial Intelligence
The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models (MLLMs) to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC.
title VIGC: Visual Instruction Generation and Correction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2308.12714