Multimodal Commonsense Knowledge Distillation for Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Shuo, Luo, Siwen, Han, Soyeon Caren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024)
by: Ding, Yihao, et al.
Published: (2024)
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024)
by: Cabral, Rina Carines, et al.
Published: (2024)
Multimodal Reranking for Knowledge-Intensive Visual Question Answering
by: Wen, Haoyang, et al.
Published: (2024)
by: Wen, Haoyang, et al.
Published: (2024)
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
by: Dai, Yue, et al.
Published: (2025)
by: Dai, Yue, et al.
Published: (2025)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
by: Zhai, Weihe, et al.
Published: (2023)
by: Zhai, Weihe, et al.
Published: (2023)
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2026)
When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoning
by: Akasaka, Muku, et al.
Published: (2026)
by: Akasaka, Muku, et al.
Published: (2026)
'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue
by: Gao, Rena, et al.
Published: (2024)
by: Gao, Rena, et al.
Published: (2024)
Declarative Knowledge Distillation from Large Language Models for Visual Question Answering Datasets
by: Eiter, Thomas, et al.
Published: (2024)
by: Eiter, Thomas, et al.
Published: (2024)
Reporting and Analysing the Environmental Impact of Language Models on the Example of Commonsense Question Answering with External Knowledge
by: Usmanova, Aida, et al.
Published: (2024)
by: Usmanova, Aida, et al.
Published: (2024)
Rationale-guided Prompting for Knowledge-based Visual Question Answering
by: Hu, Zhongjian, et al.
Published: (2024)
by: Hu, Zhongjian, et al.
Published: (2024)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering
by: Hu, Zhongjian, et al.
Published: (2024)
by: Hu, Zhongjian, et al.
Published: (2024)
Local Interpretations for Explainable Natural Language Processing: A Survey
by: Luo, Siwen, et al.
Published: (2021)
by: Luo, Siwen, et al.
Published: (2021)
CoDi: Conversational Distillation for Grounded Question Answering
by: Huber, Patrick, et al.
Published: (2024)
by: Huber, Patrick, et al.
Published: (2024)
Self-Correction Distillation for Structured Data Question Answering
by: Zhu, Yushan, et al.
Published: (2025)
by: Zhu, Yushan, et al.
Published: (2025)
Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
by: Pan, Jiqun, et al.
Published: (2025)
by: Pan, Jiqun, et al.
Published: (2025)
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering
by: Cao, Han, et al.
Published: (2024)
by: Cao, Han, et al.
Published: (2024)
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
by: Kim, Wonjoong, et al.
Published: (2024)
by: Kim, Wonjoong, et al.
Published: (2024)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
by: Cocchi, Federico, et al.
Published: (2024)
by: Cocchi, Federico, et al.
Published: (2024)
Efficient Medical Question Answering with Knowledge-Augmented Question Generation
by: Khlaut, Julien, et al.
Published: (2024)
by: Khlaut, Julien, et al.
Published: (2024)
What Really is Commonsense Knowledge?
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
Location-Aware Pretraining for Medical Difference Visual Question Answering
by: Musinguzi, Denis, et al.
Published: (2026)
by: Musinguzi, Denis, et al.
Published: (2026)
Subgraph Retrieval Enhanced by Graph-Text Alignment for Commonsense Question Answering
by: Peng, Boci, et al.
Published: (2024)
by: Peng, Boci, et al.
Published: (2024)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Multimodal Multihop Source Retrieval for Web Question Answering
by: Yarrabelly, Navya, et al.
Published: (2025)
by: Yarrabelly, Navya, et al.
Published: (2025)
Conversational Question Answering with Reformulations over Knowledge Graph
by: Liu, Lihui, et al.
Published: (2023)
by: Liu, Lihui, et al.
Published: (2023)
iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question Answering
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Question-to-Question Retrieval for Hallucination-Free Knowledge Access: An Approach for Wikipedia and Wikidata Question Answering
by: Thottingal, Santhosh
Published: (2025)
by: Thottingal, Santhosh
Published: (2025)
MKRAG: Medical Knowledge Retrieval Augmented Generation for Medical Question Answering
by: Shi, Yucheng, et al.
Published: (2023)
by: Shi, Yucheng, et al.
Published: (2023)
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning
by: Palta, Shramay, et al.
Published: (2024)
by: Palta, Shramay, et al.
Published: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
LinkQ: An LLM-Assisted Visual Interface for Knowledge Graph Question-Answering
by: Li, Harry, et al.
Published: (2024)
by: Li, Harry, et al.
Published: (2024)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
by: Li, Shuo, et al.
Published: (2023)
by: Li, Shuo, et al.
Published: (2023)
Interpretable Question Answering with Knowledge Graphs
by: Aneja, Kartikeya, et al.
Published: (2025)
by: Aneja, Kartikeya, et al.
Published: (2025)
Reinforcement Learning for Conversational Question Answering over Knowledge Graph
by: Wu, Mi
Published: (2024)
by: Wu, Mi
Published: (2024)
Dynamic Few-Shot Learning for Knowledge Graph Question Answering
by: D'Abramo, Jacopo, et al.
Published: (2024)
by: D'Abramo, Jacopo, et al.
Published: (2024)
Similar Items
-
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering
by: Yang, Shuo, et al.
Published: (2025) -
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering
by: Ding, Yihao, et al.
Published: (2024) -
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling
by: Wang, Eileen, et al.
Published: (2024) -
3M-Health: Multimodal Multi-Teacher Knowledge Distillation for Mental Health Detection
by: Cabral, Rina Carines, et al.
Published: (2024) -
Multimodal Reranking for Knowledge-Intensive Visual Question Answering
by: Wen, Haoyang, et al.
Published: (2024)