LLaVA-Video: Video Instruction Tuning With Synthetic Data
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911087308308480 |
|---|---|
| author | Zhang, Yuanhan Wu, Jinming Li, Wei Li, Bo Ma, Zejun Liu, Ziwei Li, Chunyuan |
| author_facet | Zhang, Yuanhan Wu, Jinming Li, Wei Li, Bo Ma, Zejun Liu, Ziwei Li, Chunyuan |
| contents | The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02713 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LLaVA-Video: Video Instruction Tuning With Synthetic Data Zhang, Yuanhan Wu, Jinming Li, Wei Li, Bo Ma, Zejun Liu, Ziwei Li, Chunyuan Computer Vision and Pattern Recognition Computation and Language The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints. |
| title | LLaVA-Video: Video Instruction Tuning With Synthetic Data |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2410.02713 |