VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916439849435136 |
|---|---|
| author | Lan, Xiaohan Yuan, Yitian Jie, Zequn Ma, Lin |
| author_facet | Lan, Xiaohan Yuan, Yitian Jie, Zequn Ma, Lin |
| contents | Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these challenges, we propose VidCompress, a novel Video-LLM featuring memory-enhanced temporal compression. VidCompress employs a dual-compressor approach: a memory-enhanced compressor captures both short-term and long-term temporal relationships in videos and compresses the visual tokens using a multiscale transformer with a memory-cache mechanism, while a text-perceived compressor generates condensed visual tokens by utilizing Q-Former and integrating temporal contexts into query embeddings with cross attention. Experiments on several VideoQA datasets and comprehensive benchmarks demonstrate that VidCompress efficiently models complex temporal-spatial relations and significantly outperforms existing Video-LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_11417 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models Lan, Xiaohan Yuan, Yitian Jie, Zequn Ma, Lin Computer Vision and Pattern Recognition Multimedia Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these challenges, we propose VidCompress, a novel Video-LLM featuring memory-enhanced temporal compression. VidCompress employs a dual-compressor approach: a memory-enhanced compressor captures both short-term and long-term temporal relationships in videos and compresses the visual tokens using a multiscale transformer with a memory-cache mechanism, while a text-perceived compressor generates condensed visual tokens by utilizing Q-Former and integrating temporal contexts into query embeddings with cross attention. Experiments on several VideoQA datasets and comprehensive benchmarks demonstrate that VidCompress efficiently models complex temporal-spatial relations and significantly outperforms existing Video-LLMs. |
| title | VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2410.11417 |