Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weng, Weixi, Zhu, Jieming, Meng, Xiaojun, Zhang, Hao, Zhang, Rui, Yuan, Chun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916593106157568
author Weng, Weixi
Zhu, Jieming
Meng, Xiaojun
Zhang, Hao
Zhang, Rui
Yuan, Chun
author_facet Weng, Weixi
Zhu, Jieming
Meng, Xiaojun
Zhang, Hao
Zhang, Rui
Yuan, Chun
contents Multimodal large language models (MLLMs) have demonstrated great performance on visual question answering (VQA). When it comes to knowledge-based Visual Question Answering (KB-VQA), MLLMs may lack the specialized domain knowledge needed to answer questions, necessitating the retrieval of necessary information from external knowledge sources. Previous works like Retrival-Augmented VQA-v2 (RAVQA-v2) focus on utilizing as much input information, such as image-based textual descriptions and retrieved knowledge, as possible to improve performance, but they all overlook the issue that with the number of input tokens increasing, inference efficiency significantly decreases, which contradicts the demands of practical applications. To address this issue, we propose \textbf{R}etrieval-\textbf{A}ugmented MLLMs with Compressed Contexts (RACC). RACC learns to compress and aggregate retrieved knowledge for a given image-question pair, generating a compact modulation in the form of Key-Value (KV) cache to adapt the downstream frozen MLLM, thereby achieving effective and efficient inference. RACC achieves a state-of-the-art (SOTA) performance of 63.92\% on OK-VQA. Moreover, it significantly reduces inference latency by 22.0\%-59.7\% compared to the prominent RAVQA-v2. Abundant experiments show RACC's broad applicability. It is compatible with various off-the-shelf MLLMs and can also handle different knowledge sources including textual and multimodal documents.
format Preprint
id arxiv_https___arxiv_org_abs_2409_07331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
Weng, Weixi
Zhu, Jieming
Meng, Xiaojun
Zhang, Hao
Zhang, Rui
Yuan, Chun
Computer Vision and Pattern Recognition
Machine Learning
Multimodal large language models (MLLMs) have demonstrated great performance on visual question answering (VQA). When it comes to knowledge-based Visual Question Answering (KB-VQA), MLLMs may lack the specialized domain knowledge needed to answer questions, necessitating the retrieval of necessary information from external knowledge sources. Previous works like Retrival-Augmented VQA-v2 (RAVQA-v2) focus on utilizing as much input information, such as image-based textual descriptions and retrieved knowledge, as possible to improve performance, but they all overlook the issue that with the number of input tokens increasing, inference efficiency significantly decreases, which contradicts the demands of practical applications. To address this issue, we propose \textbf{R}etrieval-\textbf{A}ugmented MLLMs with Compressed Contexts (RACC). RACC learns to compress and aggregate retrieved knowledge for a given image-question pair, generating a compact modulation in the form of Key-Value (KV) cache to adapt the downstream frozen MLLM, thereby achieving effective and efficient inference. RACC achieves a state-of-the-art (SOTA) performance of 63.92\% on OK-VQA. Moreover, it significantly reduces inference latency by 22.0\%-59.7\% compared to the prominent RAVQA-v2. Abundant experiments show RACC's broad applicability. It is compatible with various off-the-shelf MLLMs and can also handle different knowledge sources including textual and multimodal documents.
title Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2409.07331