Hallucination Mitigation Prompts Long-term Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yiwei, Liu, Zhihang, Liu, Chuanbin, Pu, Bowei, Zhang, Zhihan, Xie, Hongtao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913393474011136
author Sun, Yiwei
Liu, Zhihang
Liu, Chuanbin
Pu, Bowei
Zhang, Zhihan
Xie, Hongtao
author_facet Sun, Yiwei
Liu, Zhihang
Liu, Chuanbin
Pu, Bowei
Zhang, Zhihan
Xie, Hongtao
contents Recently, multimodal large language models have made significant advancements in video understanding tasks. However, their ability to understand unprocessed long videos is very limited, primarily due to the difficulty in supporting the enormous memory overhead. Although existing methods achieve a balance between memory and information by aggregating frames, they inevitably introduce the severe hallucination issue. To address this issue, this paper constructs a comprehensive hallucination mitigation pipeline based on existing MLLMs. Specifically, we use the CLIP Score to guide the frame sampling process with questions, selecting key frames relevant to the question. Then, We inject question information into the queries of the image Q-former to obtain more important visual features. Finally, during the answer generation stage, we utilize chain-of-thought and in-context learning techniques to explicitly control the generation of answers. It is worth mentioning that for the breakpoint mode, we found that image understanding models achieved better results than video understanding models. Therefore, we aggregated the answers from both types of models using a comparison mechanism. Ultimately, We achieved 84.2\% and 62.9\% for the global and breakpoint modes respectively on the MovieChat dataset, surpassing the official baseline model by 29.1\% and 24.1\%. Moreover the proposed method won the third place in the CVPR LOVEU 2024 Long-Term Video Question Answering Challenge. The code is avaiable at https://github.com/lntzm/CVPR24Track-LongVideo
format Preprint
id arxiv_https___arxiv_org_abs_2406_11333
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hallucination Mitigation Prompts Long-term Video Understanding
Sun, Yiwei
Liu, Zhihang
Liu, Chuanbin
Pu, Bowei
Zhang, Zhihan
Xie, Hongtao
Computer Vision and Pattern Recognition
Recently, multimodal large language models have made significant advancements in video understanding tasks. However, their ability to understand unprocessed long videos is very limited, primarily due to the difficulty in supporting the enormous memory overhead. Although existing methods achieve a balance between memory and information by aggregating frames, they inevitably introduce the severe hallucination issue. To address this issue, this paper constructs a comprehensive hallucination mitigation pipeline based on existing MLLMs. Specifically, we use the CLIP Score to guide the frame sampling process with questions, selecting key frames relevant to the question. Then, We inject question information into the queries of the image Q-former to obtain more important visual features. Finally, during the answer generation stage, we utilize chain-of-thought and in-context learning techniques to explicitly control the generation of answers. It is worth mentioning that for the breakpoint mode, we found that image understanding models achieved better results than video understanding models. Therefore, we aggregated the answers from both types of models using a comparison mechanism. Ultimately, We achieved 84.2\% and 62.9\% for the global and breakpoint modes respectively on the MovieChat dataset, surpassing the official baseline model by 29.1\% and 24.1\%. Moreover the proposed method won the third place in the CVPR LOVEU 2024 Long-Term Video Question Answering Challenge. The code is avaiable at https://github.com/lntzm/CVPR24Track-LongVideo
title Hallucination Mitigation Prompts Long-term Video Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.11333