Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thai, Triet Minh, Luu, Son T.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866907942133956608
author Thai, Triet Minh
Luu, Son T.
author_facet Thai, Triet Minh
Luu, Son T.
contents Visual Question Answering (VQA) is a task that requires computers to give correct answers for the input questions based on the images. This task can be solved by humans with ease but is a challenge for computers. The VLSP2022-EVJVQA shared task carries the Visual Question Answering task in the multilingual domain on a newly released dataset: UIT-EVJVQA, in which the questions and answers are written in three different languages: English, Vietnamese and Japanese. We approached the challenge as a sequence-to-sequence learning task, in which we integrated hints from pre-trained state-of-the-art VQA models and image features with Convolutional Sequence-to-Sequence network to generate the desired answers. Our results obtained up to 0.3442 by F1 score on the public test set, 0.4210 on the private test set, and placed 3rd in the competition.
format Preprint
id arxiv_https___arxiv_org_abs_2303_12671
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
Thai, Triet Minh
Luu, Son T.
Computer Vision and Pattern Recognition
Computation and Language
Visual Question Answering (VQA) is a task that requires computers to give correct answers for the input questions based on the images. This task can be solved by humans with ease but is a challenge for computers. The VLSP2022-EVJVQA shared task carries the Visual Question Answering task in the multilingual domain on a newly released dataset: UIT-EVJVQA, in which the questions and answers are written in three different languages: English, Vietnamese and Japanese. We approached the challenge as a sequence-to-sequence learning task, in which we integrated hints from pre-trained state-of-the-art VQA models and image features with Convolutional Sequence-to-Sequence network to generate the desired answers. Our results obtained up to 0.3442 by F1 score on the public test set, 0.4210 on the private test set, and placed 3rd in the competition.
title Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2303.12671