Improved Baselines with Visual Instruction Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Haotian, Li, Chunyuan, Li, Yuheng, Lee, Yong Jae
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911877560270848
author Liu, Haotian
Li, Chunyuan
Li, Yuheng
Lee, Yong Jae
author_facet Liu, Haotian
Li, Chunyuan
Li, Yuheng
Lee, Yong Jae
contents Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerful and data-efficient. With simple modifications to LLaVA, namely, using CLIP-ViT-L-336px with an MLP projection and adding academic-task-oriented VQA data with simple response formatting prompts, we establish stronger baselines that achieve state-of-the-art across 11 benchmarks. Our final 13B checkpoint uses merely 1.2M publicly available data, and finishes full training in ~1 day on a single 8-A100 node. We hope this can make state-of-the-art LMM research more accessible. Code and model will be publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03744
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improved Baselines with Visual Instruction Tuning
Liu, Haotian
Li, Chunyuan
Li, Yuheng
Lee, Yong Jae
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerful and data-efficient. With simple modifications to LLaVA, namely, using CLIP-ViT-L-336px with an MLP projection and adding academic-task-oriented VQA data with simple response formatting prompts, we establish stronger baselines that achieve state-of-the-art across 11 benchmarks. Our final 13B checkpoint uses merely 1.2M publicly available data, and finishes full training in ~1 day on a single 8-A100 node. We hope this can make state-of-the-art LMM research more accessible. Code and model will be publicly available.
title Improved Baselines with Visual Instruction Tuning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.03744