A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Shijie, Zhang, Ruiyi, Zhou, Yufan, Chen, Changyou
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913622621421568
author Zhou, Shijie
Zhang, Ruiyi
Zhou, Yufan
Chen, Changyou
author_facet Zhou, Shijie
Zhang, Ruiyi
Zhou, Yufan
Chen, Changyou
contents Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but its quality is poor, as multimodal alignment remains a hurdle even for the largest models. In this work, we propose LLaVAR-2, to enhance multimodal alignment for text-rich images through hybrid instruction generation between human annotators and large language models. Specifically, it involves detailed image captions from human annotators, followed by the use of these annotations in tailored text prompts for GPT-4o to curate a dataset. It also implements several mechanisms to filter out low-quality data, and the resulting dataset comprises 424k high-quality pairs of instructions. Empirical results show that models fine-tuned on this dataset exhibit impressive enhancements over those trained with self-instruct data.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16364
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
Zhou, Shijie
Zhang, Ruiyi
Zhou, Yufan
Chen, Changyou
Computer Vision and Pattern Recognition
Computation and Language
Large multimodal models still struggle with text-rich images because of inadequate training data. Self-Instruct provides an annotation-free way for generating instruction data, but its quality is poor, as multimodal alignment remains a hurdle even for the largest models. In this work, we propose LLaVAR-2, to enhance multimodal alignment for text-rich images through hybrid instruction generation between human annotators and large language models. Specifically, it involves detailed image captions from human annotators, followed by the use of these annotations in tailored text prompts for GPT-4o to curate a dataset. It also implements several mechanisms to filter out low-quality data, and the resulting dataset comprises 424k high-quality pairs of instructions. Empirical results show that models fine-tuned on this dataset exhibit impressive enhancements over those trained with self-instruct data.
title A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.16364