Saved in:
Bibliographic Details
Main Authors: Kuang, Jiayi, Xie, Jingyou, Luo, Haohao, Li, Ronghao, Xu, Zhe, Cheng, Xianfeng, Li, Yinghui, Lin, Xika, Shen, Ying
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.17558
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917849191153664
author Kuang, Jiayi
Xie, Jingyou
Luo, Haohao
Li, Ronghao
Xu, Zhe
Cheng, Xianfeng
Li, Yinghui
Lin, Xika
Shen, Ying
author_facet Kuang, Jiayi
Xie, Jingyou
Luo, Haohao
Li, Ronghao
Xu, Zhe
Cheng, Xianfeng
Li, Yinghui
Lin, Xika
Shen, Ying
contents Visual Question Answering (VQA) is a challenge task that combines natural language processing and computer vision techniques and gradually becomes a benchmark test task in multimodal large language models (MLLMs). The goal of our survey is to provide an overview of the development of VQA and a detailed description of the latest models with high timeliness. This survey gives an up-to-date synthesis of natural language understanding of images and text, as well as the knowledge reasoning module based on image-question information on the core VQA tasks. In addition, we elaborate on recent advances in extracting and fusing modal information with vision-language pretraining models and multimodal large language models in VQA. We also exhaustively review the progress of knowledge reasoning in VQA by detailing the extraction of internal knowledge and the introduction of external knowledge. Finally, we present the datasets of VQA and different evaluation metrics and discuss possible directions for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17558
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
Kuang, Jiayi
Xie, Jingyou
Luo, Haohao
Li, Ronghao
Xu, Zhe
Cheng, Xianfeng
Li, Yinghui
Lin, Xika
Shen, Ying
Computation and Language
Computer Vision and Pattern Recognition
Visual Question Answering (VQA) is a challenge task that combines natural language processing and computer vision techniques and gradually becomes a benchmark test task in multimodal large language models (MLLMs). The goal of our survey is to provide an overview of the development of VQA and a detailed description of the latest models with high timeliness. This survey gives an up-to-date synthesis of natural language understanding of images and text, as well as the knowledge reasoning module based on image-question information on the core VQA tasks. In addition, we elaborate on recent advances in extracting and fusing modal information with vision-language pretraining models and multimodal large language models in VQA. We also exhaustively review the progress of knowledge reasoning in VQA by detailing the extraction of internal knowledge and the introduction of external knowledge. Finally, we present the datasets of VQA and different evaluation metrics and discuss possible directions for future work.
title Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.17558