Visual Grounding Methods for VQA are Working for the Wrong Reasons!

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shrestha, Robik, Kafle, Kushal, Kanan, Christopher
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910418868371456
author Shrestha, Robik
Kafle, Kushal
Kanan, Christopher
author_facet Shrestha, Robik
Kafle, Kushal
Kanan, Christopher
contents Existing Visual Question Answering (VQA) methods tend to exploit dataset biases and spurious statistical correlations, instead of producing right answers for the right reasons. To address this issue, recent bias mitigation methods for VQA propose to incorporate visual cues (e.g., human attention maps) to better ground the VQA models, showcasing impressive gains. However, we show that the performance improvements are not a result of improved visual grounding, but a regularization effect which prevents over-fitting to linguistic priors. For instance, we find that it is not actually necessary to provide proper, human-based cues; random, insensible cues also result in similar improvements. Based on this observation, we propose a simpler regularization scheme that does not require any external annotations and yet achieves near state-of-the-art performance on VQA-CPv2.
format Preprint
id arxiv_https___arxiv_org_abs_2004_05704
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Visual Grounding Methods for VQA are Working for the Wrong Reasons!
Shrestha, Robik
Kafle, Kushal
Kanan, Christopher
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Existing Visual Question Answering (VQA) methods tend to exploit dataset biases and spurious statistical correlations, instead of producing right answers for the right reasons. To address this issue, recent bias mitigation methods for VQA propose to incorporate visual cues (e.g., human attention maps) to better ground the VQA models, showcasing impressive gains. However, we show that the performance improvements are not a result of improved visual grounding, but a regularization effect which prevents over-fitting to linguistic priors. For instance, we find that it is not actually necessary to provide proper, human-based cues; random, insensible cues also result in similar improvements. Based on this observation, we propose a simpler regularization scheme that does not require any external annotations and yet achieves near state-of-the-art performance on VQA-CPv2.
title Visual Grounding Methods for VQA are Working for the Wrong Reasons!
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2004.05704