Multimodal Integration of Human-Like Attention in Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sood, Ekta, Kögel, Fabian, Müller, Philipp, Thomas, Dominike, Bace, Mihai, Bulling, Andreas
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910038261497856
author Sood, Ekta
Kögel, Fabian
Müller, Philipp
Thomas, Dominike
Bace, Mihai
Bulling, Andreas
author_facet Sood, Ekta
Kögel, Fabian
Müller, Philipp
Thomas, Dominike
Bace, Mihai
Bulling, Andreas
contents Human-like attention as a supervisory signal to guide neural attention has shown significant promise but is currently limited to uni-modal integration - even for inherently multimodal tasks such as visual question answering (VQA). We present the Multimodal Human-like Attention Network (MULAN) - the first method for multimodal integration of human-like attention on image and text during training of VQA models. MULAN integrates attention predictions from two state-of-the-art text and image saliency models into neural self-attention layers of a recent transformer-based VQA model. Through evaluations on the challenging VQAv2 dataset, we show that MULAN achieves a new state-of-the-art performance of 73.98% accuracy on test-std and 73.72% on test-dev and, at the same time, has approximately 80% fewer trainable parameters than prior work. Overall, our work underlines the potential of integrating multimodal human-like and neural attention for VQA
format Preprint
id arxiv_https___arxiv_org_abs_2109_13139
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Multimodal Integration of Human-Like Attention in Visual Question Answering
Sood, Ekta
Kögel, Fabian
Müller, Philipp
Thomas, Dominike
Bace, Mihai
Bulling, Andreas
Computer Vision and Pattern Recognition
Computation and Language
Human-like attention as a supervisory signal to guide neural attention has shown significant promise but is currently limited to uni-modal integration - even for inherently multimodal tasks such as visual question answering (VQA). We present the Multimodal Human-like Attention Network (MULAN) - the first method for multimodal integration of human-like attention on image and text during training of VQA models. MULAN integrates attention predictions from two state-of-the-art text and image saliency models into neural self-attention layers of a recent transformer-based VQA model. Through evaluations on the challenging VQAv2 dataset, we show that MULAN achieves a new state-of-the-art performance of 73.98% accuracy on test-std and 73.72% on test-dev and, at the same time, has approximately 80% fewer trainable parameters than prior work. Overall, our work underlines the potential of integrating multimodal human-like and neural attention for VQA
title Multimodal Integration of Human-Like Attention in Visual Question Answering
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2109.13139