Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915061417639936 |
|---|---|
| author | Wang, Han Nie, Yuxiang Ye, Yongjie GuanYu, Deng Wang, Yanjie Li, Shuai Yu, Haiyang Lu, Jinghui Huang, Can |
| author_facet | Wang, Han Nie, Yuxiang Ye, Yongjie GuanYu, Deng Wang, Yanjie Li, Shuai Yu, Haiyang Lu, Jinghui Huang, Can |
| contents | The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_09530 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM Wang, Han Nie, Yuxiang Ye, Yongjie GuanYu, Deng Wang, Yanjie Li, Shuai Yu, Haiyang Lu, Jinghui Huang, Can Computer Vision and Pattern Recognition The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM |
| title | Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.09530 |