Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Han, Nie, Yuxiang, Ye, Yongjie, GuanYu, Deng, Wang, Yanjie, Li, Shuai, Yu, Haiyang, Lu, Jinghui, Huang, Can
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915061417639936
author Wang, Han
Nie, Yuxiang
Ye, Yongjie
GuanYu, Deng
Wang, Yanjie
Li, Shuai
Yu, Haiyang
Lu, Jinghui
Huang, Can
author_facet Wang, Han
Nie, Yuxiang
Ye, Yongjie
GuanYu, Deng
Wang, Yanjie
Li, Shuai
Yu, Haiyang
Lu, Jinghui
Huang, Can
contents The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM
format Preprint
id arxiv_https___arxiv_org_abs_2412_09530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Wang, Han
Nie, Yuxiang
Ye, Yongjie
GuanYu, Deng
Wang, Yanjie
Li, Shuai
Yu, Haiyang
Lu, Jinghui
Huang, Can
Computer Vision and Pattern Recognition
The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM
title Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.09530