Multimodal Commonsense Knowledge Distillation for Visual Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Shuo, Luo, Siwen, Han, Soyeon Caren
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909377780252672
author Yang, Shuo
Luo, Siwen
Han, Soyeon Caren
author_facet Yang, Shuo
Luo, Siwen
Han, Soyeon Caren
contents Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). However, these models struggle with VQA questions that require external commonsense knowledge due to the challenges in generating high-quality prompts and the high computational costs of fine-tuning. In this work, we propose a novel graph-based multimodal commonsense knowledge distillation framework that constructs a unified relational graph over commonsense knowledge, visual objects and questions through a Graph Convolutional Network (GCN) following a teacher-student environment. This proposed framework is flexible with any type of teacher and student models without further fine-tuning, and has achieved competitive performances on the ScienceQA dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Commonsense Knowledge Distillation for Visual Question Answering
Yang, Shuo
Luo, Siwen
Han, Soyeon Caren
Computation and Language
Artificial Intelligence
Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). However, these models struggle with VQA questions that require external commonsense knowledge due to the challenges in generating high-quality prompts and the high computational costs of fine-tuning. In this work, we propose a novel graph-based multimodal commonsense knowledge distillation framework that constructs a unified relational graph over commonsense knowledge, visual objects and questions through a Graph Convolutional Network (GCN) following a teacher-student environment. This proposed framework is flexible with any type of teacher and student models without further fine-tuning, and has achieved competitive performances on the ScienceQA dataset.
title Multimodal Commonsense Knowledge Distillation for Visual Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.02722