LLaVA-Video: Video Instruction Tuning With Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuanhan, Wu, Jinming, Li, Wei, Li, Bo, Ma, Zejun, Liu, Ziwei, Li, Chunyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911087308308480
author Zhang, Yuanhan
Wu, Jinming
Li, Wei
Li, Bo
Ma, Zejun
Liu, Ziwei
Li, Chunyuan
author_facet Zhang, Yuanhan
Wu, Jinming
Li, Wei
Li, Bo
Ma, Zejun
Liu, Ziwei
Li, Chunyuan
contents The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02713
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaVA-Video: Video Instruction Tuning With Synthetic Data
Zhang, Yuanhan
Wu, Jinming
Li, Wei
Li, Bo
Ma, Zejun
Liu, Ziwei
Li, Chunyuan
Computer Vision and Pattern Recognition
Computation and Language
The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints.
title LLaVA-Video: Video Instruction Tuning With Synthetic Data
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2410.02713