Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ziyan, Kafle, Kushal, Dernoncourt, Franck, Ordonez, Vicente
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929200117579776
author Yang, Ziyan
Kafle, Kushal
Dernoncourt, Franck
Ordonez, Vicente
author_facet Yang, Ziyan
Kafle, Kushal
Dernoncourt, Franck
Ordonez, Vicente
contents We propose a margin-based loss for tuning joint vision-language models so that their gradient-based explanations are consistent with region-level annotations provided by humans for relatively smaller grounding datasets. We refer to this objective as Attention Mask Consistency (AMC) and demonstrate that it produces superior visual grounding results than previous methods that rely on using vision-language models to score the outputs of object detectors. Particularly, a model trained with AMC on top of standard vision-language modeling objectives obtains a state-of-the-art accuracy of 86.49% in the Flickr30k visual grounding benchmark, an absolute improvement of 5.38% when compared to the best previous model trained under the same level of supervision. Our approach also performs exceedingly well on established benchmarks for referring expression comprehension where it obtains 80.34% accuracy in the easy test of RefCOCO+, and 64.55% in the difficult split. AMC is effective, easy to implement, and is general as it can be adopted by any vision-language model, and can use any type of region annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2206_15462
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
Yang, Ziyan
Kafle, Kushal
Dernoncourt, Franck
Ordonez, Vicente
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
We propose a margin-based loss for tuning joint vision-language models so that their gradient-based explanations are consistent with region-level annotations provided by humans for relatively smaller grounding datasets. We refer to this objective as Attention Mask Consistency (AMC) and demonstrate that it produces superior visual grounding results than previous methods that rely on using vision-language models to score the outputs of object detectors. Particularly, a model trained with AMC on top of standard vision-language modeling objectives obtains a state-of-the-art accuracy of 86.49% in the Flickr30k visual grounding benchmark, an absolute improvement of 5.38% when compared to the best previous model trained under the same level of supervision. Our approach also performs exceedingly well on established benchmarks for referring expression comprehension where it obtains 80.34% accuracy in the easy test of RefCOCO+, and 64.55% in the difficult split. AMC is effective, easy to implement, and is general as it can be adopted by any vision-language model, and can use any type of region annotations.
title Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2206.15462