Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Yuyang, Gu, Jiaqi, Yang, Qi, Fan, Lubin, Wu, Yue, Wang, Ying, Ding, Kun, Xiang, Shiming, Ye, Jieping
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915562758602752
author Hong, Yuyang
Gu, Jiaqi
Yang, Qi
Fan, Lubin
Wu, Yue
Wang, Ying
Ding, Kun
Xiang, Shiming
Ye, Jieping
author_facet Hong, Yuyang
Gu, Jiaqi
Yang, Qi
Fan, Lubin
Wu, Yue
Wang, Ying
Ding, Kun
Xiang, Shiming
Ye, Jieping
contents Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-augmented generation (RAG) achieves significant advances in this task by combining knowledge-base querying, it still struggles with the quality of multimodal queries and the relevance of retrieved results. To overcome these challenges, we propose a novel three-stage method, termed Wiki-PRF, including Processing, Retrieval and Filtering stages. The processing stage dynamically invokes visual tools to extract precise multimodal information for retrieval. The retrieval stage integrates visual and text features to achieve multimodal knowledge retrieval. The filtering stage performs relevance filtering and concentration on retrieval results. To this end, we introduce a visual language model trained with answer accuracy and format consistency as reward signals via a reinforcement learning manner. This enhances the model's reasoning, tool invocation for accurate queries, and filtering of irrelevant content. Experiments on benchmark datasets (E-VQA and InfoSeek) show significant improvements~(36.0 and 42.8) in answer quality, achieving state-of-the-art performance. Code is available at https://github.com/cqu-student/Wiki-PRF
format Preprint
id arxiv_https___arxiv_org_abs_2510_14605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Hong, Yuyang
Gu, Jiaqi
Yang, Qi
Fan, Lubin
Wu, Yue
Wang, Ying
Ding, Kun
Xiang, Shiming
Ye, Jieping
Computer Vision and Pattern Recognition
Artificial Intelligence
Knowledge-based visual question answering (KB-VQA) requires visual language models (VLMs) to integrate visual understanding with external knowledge retrieval. Although retrieval-augmented generation (RAG) achieves significant advances in this task by combining knowledge-base querying, it still struggles with the quality of multimodal queries and the relevance of retrieved results. To overcome these challenges, we propose a novel three-stage method, termed Wiki-PRF, including Processing, Retrieval and Filtering stages. The processing stage dynamically invokes visual tools to extract precise multimodal information for retrieval. The retrieval stage integrates visual and text features to achieve multimodal knowledge retrieval. The filtering stage performs relevance filtering and concentration on retrieval results. To this end, we introduce a visual language model trained with answer accuracy and format consistency as reward signals via a reinforcement learning manner. This enhances the model's reasoning, tool invocation for accurate queries, and filtering of irrelevant content. Experiments on benchmark datasets (E-VQA and InfoSeek) show significant improvements~(36.0 and 42.8) in answer quality, achieving state-of-the-art performance. Code is available at https://github.com/cqu-student/Wiki-PRF
title Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.14605