VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lan, Xiaohan, Yuan, Yitian, Jie, Zequn, Ma, Lin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916439849435136
author Lan, Xiaohan
Yuan, Yitian
Jie, Zequn
Ma, Lin
author_facet Lan, Xiaohan
Yuan, Yitian
Jie, Zequn
Ma, Lin
contents Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these challenges, we propose VidCompress, a novel Video-LLM featuring memory-enhanced temporal compression. VidCompress employs a dual-compressor approach: a memory-enhanced compressor captures both short-term and long-term temporal relationships in videos and compresses the visual tokens using a multiscale transformer with a memory-cache mechanism, while a text-perceived compressor generates condensed visual tokens by utilizing Q-Former and integrating temporal contexts into query embeddings with cross attention. Experiments on several VideoQA datasets and comprehensive benchmarks demonstrate that VidCompress efficiently models complex temporal-spatial relations and significantly outperforms existing Video-LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11417
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
Lan, Xiaohan
Yuan, Yitian
Jie, Zequn
Ma, Lin
Computer Vision and Pattern Recognition
Multimedia
Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks. However, most Video-LLMs treat videos as a sequential set of individual frames, which results in insufficient temporal-spatial interaction that hinders fine-grained comprehension and difficulty in processing longer videos due to limited visual token capacity. To address these challenges, we propose VidCompress, a novel Video-LLM featuring memory-enhanced temporal compression. VidCompress employs a dual-compressor approach: a memory-enhanced compressor captures both short-term and long-term temporal relationships in videos and compresses the visual tokens using a multiscale transformer with a memory-cache mechanism, while a text-perceived compressor generates condensed visual tokens by utilizing Q-Former and integrating temporal contexts into query embeddings with cross attention. Experiments on several VideoQA datasets and comprehensive benchmarks demonstrate that VidCompress efficiently models complex temporal-spatial relations and significantly outperforms existing Video-LLMs.
title VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2410.11417