Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Chen, Wang, Zheng, Gao, Jianjun, Liu, Wenyang, Lu, Ye, Zhang, Runzhong, Yap, Kim-Hui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917894489636864
author Cai, Chen
Wang, Zheng
Gao, Jianjun
Liu, Wenyang
Lu, Ye
Zhang, Runzhong
Yap, Kim-Hui
author_facet Cai, Chen
Wang, Zheng
Gao, Jianjun
Liu, Wenyang
Lu, Ye
Zhang, Runzhong
Yap, Kim-Hui
contents In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting. To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting. These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research. Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14\% accuracy on NExT-QA and 71.24\% accuracy on DramaQA, highlighting its practical relevance and effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00771
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
Cai, Chen
Wang, Zheng
Gao, Jianjun
Liu, Wenyang
Lu, Ye
Zhang, Runzhong
Yap, Kim-Hui
Computer Vision and Pattern Recognition
Computation and Language
In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting. To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting. These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research. Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14\% accuracy on NExT-QA and 71.24\% accuracy on DramaQA, highlighting its practical relevance and effectiveness.
title Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2410.00771