Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miao, Peihan, Su, Wei, Wang, Gaoang, Li, Xuewei, Li, Xi
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909134199193600
author Miao, Peihan
Su, Wei
Wang, Gaoang
Li, Xuewei
Li, Xi
author_facet Miao, Peihan
Su, Wei
Wang, Gaoang
Li, Xuewei
Li, Xi
contents As an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning. In addition, due to the diversity of visual scenes and the variation of linguistic expressions, some hard examples have much more abundant multi-grained information than others. How to aggregate multi-grained information from different modalities and extract abundant knowledge from hard examples is crucial in the REC task. To address aforementioned challenges, in this paper, we propose a Self-paced Multi-grained Cross-modal Interaction Modeling framework, which improves the language-to-vision localization ability through innovations in network structure and learning mechanism. Concretely, we design a transformer-based multi-grained cross-modal attention, which effectively utilizes the inherent multi-grained information in visual and linguistic encoders. Furthermore, considering the large variance of samples, we propose a self-paced sample informativeness learning to adaptively enhance the network learning for samples containing abundant multi-grained information. The proposed framework significantly outperforms state-of-the-art methods on widely used datasets, such as RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame datasets, demonstrating the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2204_09957
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
Miao, Peihan
Su, Wei
Wang, Gaoang
Li, Xuewei
Li, Xi
Computer Vision and Pattern Recognition
As an important and challenging problem in vision-language tasks, referring expression comprehension (REC) generally requires a large amount of multi-grained information of visual and linguistic modalities to realize accurate reasoning. In addition, due to the diversity of visual scenes and the variation of linguistic expressions, some hard examples have much more abundant multi-grained information than others. How to aggregate multi-grained information from different modalities and extract abundant knowledge from hard examples is crucial in the REC task. To address aforementioned challenges, in this paper, we propose a Self-paced Multi-grained Cross-modal Interaction Modeling framework, which improves the language-to-vision localization ability through innovations in network structure and learning mechanism. Concretely, we design a transformer-based multi-grained cross-modal attention, which effectively utilizes the inherent multi-grained information in visual and linguistic encoders. Furthermore, considering the large variance of samples, we propose a self-paced sample informativeness learning to adaptively enhance the network learning for samples containing abundant multi-grained information. The proposed framework significantly outperforms state-of-the-art methods on widely used datasets, such as RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame datasets, demonstrating the effectiveness of our method.
title Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2204.09957