MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Shuo, Luo, Siwen, Han, Soyeon Caren, Hovy, Eduard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)
by: Dai, Yue, et al.
Published: (2025)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
When Is Rank-1 Enough? Geometry-Guided Initialization for Parameter-Efficient Fine-Tuning
by: Zhao, Haoran, et al.
Published: (2026)
by: Zhao, Haoran, et al.
Published: (2026)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting
by: Park, Sehyuk, et al.
Published: (2025)
by: Park, Sehyuk, et al.
Published: (2025)
BRIDGE: Benchmark for multi-hop Reasoning In long multimodal Documents with Grounded Evidence
by: Xiang, Biao, et al.
Published: (2026)
by: Xiang, Biao, et al.
Published: (2026)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
by: Kim, Minjun, et al.
Published: (2024)
by: Kim, Minjun, et al.
Published: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
by: Sood, Ekta, et al.
Published: (2021)
by: Sood, Ekta, et al.
Published: (2021)
Multimodal Large Language Models and Tunings: Vision, Language, Sensors, Audio, and Beyond
by: Han, Soyeon Caren, et al.
Published: (2024)
by: Han, Soyeon Caren, et al.
Published: (2024)
Multimodal Reranking for Knowledge-Intensive Visual Question Answering
by: Wen, Haoyang, et al.
Published: (2024)
by: Wen, Haoyang, et al.
Published: (2024)
'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue
by: Gao, Rena, et al.
Published: (2024)
by: Gao, Rena, et al.
Published: (2024)
ARK-V1: An LLM-Agent for Knowledge Graph Question Answering Requiring Commonsense Reasoning
by: Klein, Jan-Felix, et al.
Published: (2025)
by: Klein, Jan-Felix, et al.
Published: (2025)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
by: Zhai, Weihe, et al.
Published: (2023)
by: Zhai, Weihe, et al.
Published: (2023)
Graph Neural Networks for Text Classification: A Survey
by: Wang, Kunze, et al.
Published: (2023)
by: Wang, Kunze, et al.
Published: (2023)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
by: Usmanova, Aida, et al.
Published: (2024)
by: Usmanova, Aida, et al.
Published: (2024)
BUCA: A Binary Classification Approach to Unsupervised Commonsense Question Answering
by: He, Jie, et al.
Published: (2023)
by: He, Jie, et al.
Published: (2023)
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
by: Li, Haoxuan, et al.
Published: (2025)
by: Li, Haoxuan, et al.
Published: (2025)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
by: Shaaban, Mai A., et al.
Published: (2025)
by: Shaaban, Mai A., et al.
Published: (2025)
Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering
by: Toroghi, Armin, et al.
Published: (2024)
by: Toroghi, Armin, et al.
Published: (2024)
3M: Multi-modal Multi-task Multi-teacher Learning for Game Event Detection
by: Ng, Thye Shan, et al.
Published: (2024)
by: Ng, Thye Shan, et al.
Published: (2024)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
by: Yeh, Yahsin, et al.
Published: (2025)
by: Yeh, Yahsin, et al.
Published: (2025)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
by: Ghosh, Shiv, et al.
Published: (2026)
by: Ghosh, Shiv, et al.
Published: (2026)
ZEBRA: Zero-Shot Example-Based Retrieval Augmentation for Commonsense Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2024)
by: Molfese, Francesco Maria, et al.
Published: (2024)
Efficient Multimodal Planning Agent for Visual Question-Answering
by: Chen, Zhuo, et al.
Published: (2026)
by: Chen, Zhuo, et al.
Published: (2026)
Knowledge-Based Counterfactual Queries for Visual Question Answering
by: Stoikou, Theodoti, et al.
Published: (2023)
by: Stoikou, Theodoti, et al.
Published: (2023)
Local Interpretations for Explainable Natural Language Processing: A Survey
by: Luo, Siwen, et al.
Published: (2021)
by: Luo, Siwen, et al.
Published: (2021)
Rationale-guided Prompting for Knowledge-based Visual Question Answering
by: Hu, Zhongjian, et al.
Published: (2024)
by: Hu, Zhongjian, et al.
Published: (2024)
Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
by: Hu, Mengkang, et al.
Published: (2024)
by: Hu, Mengkang, et al.
Published: (2024)
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
In-game Toxic Language Detection: Shared Task and Attention Residuals
by: Jia, Yuanzhe, et al.
Published: (2022)
by: Jia, Yuanzhe, et al.
Published: (2022)
Similar Items
-
Multimodal Commonsense Knowledge Distillation for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2024) -
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024) -
Deep Learning based Visually Rich Document Content Understanding: A Survey
by: Ding, Yihao, et al.
Published: (2024) -
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025) -
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)