BERT-VQA: Visual Question Answering on Plots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vu, Tai, Yang, Robert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918126727200768
author Vu, Tai
Yang, Robert
author_facet Vu, Tai
Yang, Robert
contents Visual question answering has been an exciting challenge in the field of natural language understanding, as it requires deep learning models to exchange information from both vision and language domains. In this project, we aim to tackle a subtask of this problem, namely visual question answering on plots. To achieve this, we developed BERT-VQA, a VisualBERT-based model architecture with a pretrained ResNet 101 image encoder, along with a potential addition of joint fusion. We trained and evaluated this model against a baseline that consisted of a LSTM, a CNN, and a shallow classifier. The final outcome disproved our core hypothesis that the cross-modality module in VisualBERT is essential in aligning plot components with question phrases. Therefore, our work provided valuable insights into the difficulty of the plot question answering challenge as well as the appropriateness of different model architectures in solving this problem.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BERT-VQA: Visual Question Answering on Plots
Vu, Tai
Yang, Robert
Machine Learning
Computer Vision and Pattern Recognition
Visual question answering has been an exciting challenge in the field of natural language understanding, as it requires deep learning models to exchange information from both vision and language domains. In this project, we aim to tackle a subtask of this problem, namely visual question answering on plots. To achieve this, we developed BERT-VQA, a VisualBERT-based model architecture with a pretrained ResNet 101 image encoder, along with a potential addition of joint fusion. We trained and evaluated this model against a baseline that consisted of a LSTM, a CNN, and a shallow classifier. The final outcome disproved our core hypothesis that the cross-modality module in VisualBERT is essential in aligning plot components with question phrases. Therefore, our work provided valuable insights into the difficulty of the plot question answering challenge as well as the appropriateness of different model architectures in solving this problem.
title BERT-VQA: Visual Question Answering on Plots
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.13184