Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Fan, Jin, Xiaojie, Wang, Heng, Xian, Yuchen, Feng, Jiashi, Yang, Yi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909521734008832
author Ma, Fan
Jin, Xiaojie
Wang, Heng
Xian, Yuchen
Feng, Jiashi
Yang, Yi
author_facet Ma, Fan
Jin, Xiaojie
Wang, Heng
Xian, Yuchen
Feng, Jiashi
Yang, Yi
contents Recent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and employ large language models for multi-modal tasks. However, this method often leads to the generation of irrelevant content, commonly known as "hallucination", as the length of the text increases and the impact of the video diminishes. To address this problem, we propose Vista-LLaMA, a novel framework that maintains the consistent distance between all visual tokens and any language tokens, irrespective of the generated text length. Vista-LLaMA omits relative position encoding when determining attention weights between visual and text tokens, retaining the position encoding for text and text tokens. This amplifies the effect of visual tokens on text generation, especially when the relative distance is longer between visual and text tokens. The proposed attention mechanism significantly reduces the chance of producing irrelevant text related to the video content. Furthermore, we present a sequential visual projector that projects the current video frame into tokens of language space with the assistance of the previous frame. This approach not only captures the temporal relationship within the video, but also allows less visual tokens to encompass the entire video. Our approach significantly outperforms various previous methods (e.g., Video-ChatGPT, MovieChat) on four challenging open-ended video question answering benchmarks. We reach an accuracy of 60.7 on the zero-shot NExT-QA and 60.5 on the zero-shot MSRVTT-QA, setting a new state-of-the-art performance. This project is available at https://jinxxian.github.io/Vista-LLaMA.
format Preprint
id arxiv_https___arxiv_org_abs_2312_08870
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Ma, Fan
Jin, Xiaojie
Wang, Heng
Xian, Yuchen
Feng, Jiashi
Yang, Yi
Computer Vision and Pattern Recognition
Recent advances in large video-language models have displayed promising outcomes in video comprehension. Current approaches straightforwardly convert video into language tokens and employ large language models for multi-modal tasks. However, this method often leads to the generation of irrelevant content, commonly known as "hallucination", as the length of the text increases and the impact of the video diminishes. To address this problem, we propose Vista-LLaMA, a novel framework that maintains the consistent distance between all visual tokens and any language tokens, irrespective of the generated text length. Vista-LLaMA omits relative position encoding when determining attention weights between visual and text tokens, retaining the position encoding for text and text tokens. This amplifies the effect of visual tokens on text generation, especially when the relative distance is longer between visual and text tokens. The proposed attention mechanism significantly reduces the chance of producing irrelevant text related to the video content. Furthermore, we present a sequential visual projector that projects the current video frame into tokens of language space with the assistance of the previous frame. This approach not only captures the temporal relationship within the video, but also allows less visual tokens to encompass the entire video. Our approach significantly outperforms various previous methods (e.g., Video-ChatGPT, MovieChat) on four challenging open-ended video question answering benchmarks. We reach an accuracy of 60.7 on the zero-shot NExT-QA and 60.5 on the zero-shot MSRVTT-QA, setting a new state-of-the-art performance. This project is available at https://jinxxian.github.io/Vista-LLaMA.
title Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.08870