Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Kuang, Jiayi, Xie, Jingyou, Luo, Haohao, Li, Ronghao, Xu, Zhe, Cheng, Xianfeng, Li, Yinghui, Lin, Xika, Shen, Ying |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
by: Xie, Jingyou, et al.
Published: (2024)
by: Xie, Jingyou, et al.
Published: (2024)
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
by: Zhong, Liangyu, et al.
Published: (2025)
by: Zhong, Liangyu, et al.
Published: (2025)
Understanding Network Behaviors through Natural Language Question-Answering
by: Xing, Mingzhe, et al.
Published: (2025)
by: Xing, Mingzhe, et al.
Published: (2025)
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
DriveLM: Driving with Graph Visual Question Answering
by: Sima, Chonghao, et al.
Published: (2023)
by: Sima, Chonghao, et al.
Published: (2023)
Dynamic Demonstration Retrieval and Cognitive Understanding for Emotional Support Conversation
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
by: Wang, Zining, et al.
Published: (2025)
by: Wang, Zining, et al.
Published: (2025)
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering
by: Lim, Su Hyeon, et al.
Published: (2024)
by: Lim, Su Hyeon, et al.
Published: (2024)
Reconstruction as a Bridge for Event-Based Visual Question Answering
by: Lou, Hanyue, et al.
Published: (2025)
by: Lou, Hanyue, et al.
Published: (2025)
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
by: Li, Yangyi, et al.
Published: (2025)
by: Li, Yangyi, et al.
Published: (2025)
IIU: Independent Inference Units for Knowledge-based Visual Question Answering
by: Li, Yili, et al.
Published: (2024)
by: Li, Yili, et al.
Published: (2024)
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
by: Ding, Yihao, et al.
Published: (2025)
by: Ding, Yihao, et al.
Published: (2025)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
by: Jiao, Qirui, et al.
Published: (2025)
by: Jiao, Qirui, et al.
Published: (2025)
SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation
by: Ramirez, David F., et al.
Published: (2026)
by: Ramirez, David F., et al.
Published: (2026)
A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
by: Qiu, Weikang, et al.
Published: (2025)
by: Qiu, Weikang, et al.
Published: (2025)
Visually Interpretable Subtask Reasoning for Visual Question Answering
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering
by: Zhao, Yu, et al.
Published: (2025)
by: Zhao, Yu, et al.
Published: (2025)
An In-Context Schema Understanding Method for Knowledge Base Question Answering
by: Liu, Yantao, et al.
Published: (2023)
by: Liu, Yantao, et al.
Published: (2023)
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention
by: Liu, Ying, et al.
Published: (2024)
by: Liu, Ying, et al.
Published: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
Sports Intelligence: Assessing the Sports Understanding Capabilities of Language Models through Question Answering from Text to Video
by: Yang, Zhengbang, et al.
Published: (2024)
by: Yang, Zhengbang, et al.
Published: (2024)
A Survey of Large Language Model Agents for Question Answering
by: Yue, Murong
Published: (2025)
by: Yue, Murong
Published: (2025)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
by: Kabir, Raihan, et al.
Published: (2024)
by: Kabir, Raihan, et al.
Published: (2024)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
by: Maryam, Hiba, et al.
Published: (2024)
by: Maryam, Hiba, et al.
Published: (2024)
Exploring Diverse Methods in Visual Question Answering
by: Li, Panfeng, et al.
Published: (2024)
by: Li, Panfeng, et al.
Published: (2024)
DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding
by: Tao, Mingzhe, et al.
Published: (2026)
by: Tao, Mingzhe, et al.
Published: (2026)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
Biomedical Question Answering: A Survey of Approaches and Challenges
by: Jin, Qiao, et al.
Published: (2021)
by: Jin, Qiao, et al.
Published: (2021)
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024)
by: Kan, Shichao, et al.
Published: (2024)
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
by: Zhu, Yingjian, et al.
Published: (2026)
by: Zhu, Yingjian, et al.
Published: (2026)
Efficient Multimodal Planning Agent for Visual Question-Answering
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
Fine-tuning Large Language Models for Improving Factuality in Legal Question Answering
by: Hu, Yinghao, et al.
Published: (2025)
by: Hu, Yinghao, et al.
Published: (2025)
Similar Items
-
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
by: Xie, Jingyou, et al.
Published: (2024) -
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
by: Zhong, Liangyu, et al.
Published: (2025) -
Understanding Network Behaviors through Natural Language Question-Answering
by: Xing, Mingzhe, et al.
Published: (2025) -
Multimodal Rationales for Explainable Visual Question Answering
by: Li, Kun, et al.
Published: (2024) -
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)