SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Jungang, Tao, Sicheng, Yan, Yibo, Gu, Xiaojie, Xu, Haodong, Zheng, Xu, Lyu, Yuanhuiyi, Zhang, Linfeng, Hu, Xuming
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913607164362752
author Li, Jungang
Tao, Sicheng
Yan, Yibo
Gu, Xiaojie
Xu, Haodong
Zheng, Xu
Lyu, Yuanhuiyi
Zhang, Linfeng
Hu, Xuming
author_facet Li, Jungang
Tao, Sicheng
Yan, Yibo
Gu, Xiaojie
Xu, Haodong
Zheng, Xu
Lyu, Yuanhuiyi
Zhang, Linfeng
Hu, Xuming
contents Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich and diverse audio-visual information inherent in long videos, which is crucial for comprehensive understanding. This raises the question: how can we leverage embedded audio-visual information to enhance long video understanding? Therefore, (i) we introduce SAVEn-Vid, the first-ever long audio-visual video dataset comprising over 58k audio-visual instructions. (ii) From the model perspective, we propose a time-aware Audio-Visual Large Language Model (AV-LLM), SAVEnVideo, fine-tuned on SAVEn-Vid. (iii) Besides, we present AVBench, a benchmark containing 2,500 QAs designed to evaluate models on enhanced audio-visual comprehension tasks within long video, challenging their ability to handle intricate audio-visual interactions. Experiments on AVBench reveal the limitations of current AV-LLMs. Experiments also demonstrate that SAVEnVideo outperforms the best Video-LLM by 3.61% on the zero-shot long video task (Video-MME) and surpasses the leading audio-visual LLM by 1.29% on the zero-shot audio-visual task (Music-AVQA). Consequently, at the 7B parameter scale, SAVEnVideo can achieve state-of-the-art performance. Our dataset and code will be released at https://ljungang.github.io/SAVEn-Vid/ upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
Li, Jungang
Tao, Sicheng
Yan, Yibo
Gu, Xiaojie
Xu, Haodong
Zheng, Xu
Lyu, Yuanhuiyi
Zhang, Linfeng
Hu, Xuming
Computer Vision and Pattern Recognition
Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich and diverse audio-visual information inherent in long videos, which is crucial for comprehensive understanding. This raises the question: how can we leverage embedded audio-visual information to enhance long video understanding? Therefore, (i) we introduce SAVEn-Vid, the first-ever long audio-visual video dataset comprising over 58k audio-visual instructions. (ii) From the model perspective, we propose a time-aware Audio-Visual Large Language Model (AV-LLM), SAVEnVideo, fine-tuned on SAVEn-Vid. (iii) Besides, we present AVBench, a benchmark containing 2,500 QAs designed to evaluate models on enhanced audio-visual comprehension tasks within long video, challenging their ability to handle intricate audio-visual interactions. Experiments on AVBench reveal the limitations of current AV-LLMs. Experiments also demonstrate that SAVEnVideo outperforms the best Video-LLM by 3.61% on the zero-shot long video task (Video-MME) and surpasses the leading audio-visual LLM by 1.29% on the zero-shot audio-visual task (Music-AVQA). Consequently, at the 7B parameter scale, SAVEnVideo can achieve state-of-the-art performance. Our dataset and code will be released at https://ljungang.github.io/SAVEn-Vid/ upon acceptance.
title SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16213