Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hui, Tianrui, Ding, Zihan, Huang, Junshi, Wei, Xiaoming, Wei, Xiaolin, Dai, Jiao, Han, Jizhong, Liu, Si
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911792189407232
author Hui, Tianrui
Ding, Zihan
Huang, Junshi
Wei, Xiaoming
Wei, Xiaolin
Dai, Jiao
Han, Jizhong
Liu, Si
author_facet Hui, Tianrui
Ding, Zihan
Huang, Junshi
Wei, Xiaoming
Wei, Xiaolin
Dai, Jiao
Han, Jizhong
Liu, Si
contents Panoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual contexts from offline-generated mask proposals to phrase features, which tend to be noisy and fragmentary. The recent one-stage method aggregates only pixel contexts from image features to phrase features, which may incur semantic misalignment due to lacking object priors. To realize more comprehensive visual-linguistic interaction, we propose to enrich phrases with coupled pixel and object contexts by designing a Phrase-Pixel-Object Transformer Decoder (PPO-TD), where both fine-grained part details and coarse-grained entity clues are aggregated to phrase features. In addition, we also propose a PhraseObject Contrastive Loss (POCL) to pull closer the matched phrase-object pairs and push away unmatched ones for aggregating more precise object contexts from more phrase-relevant object tokens. Extensive experiments on the PNG benchmark show our method achieves new state-of-the-art performance with large margins.
format Preprint
id arxiv_https___arxiv_org_abs_2311_01091
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
Hui, Tianrui
Ding, Zihan
Huang, Junshi
Wei, Xiaoming
Wei, Xiaolin
Dai, Jiao
Han, Jizhong
Liu, Si
Computer Vision and Pattern Recognition
Panoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual contexts from offline-generated mask proposals to phrase features, which tend to be noisy and fragmentary. The recent one-stage method aggregates only pixel contexts from image features to phrase features, which may incur semantic misalignment due to lacking object priors. To realize more comprehensive visual-linguistic interaction, we propose to enrich phrases with coupled pixel and object contexts by designing a Phrase-Pixel-Object Transformer Decoder (PPO-TD), where both fine-grained part details and coarse-grained entity clues are aggregated to phrase features. In addition, we also propose a PhraseObject Contrastive Loss (POCL) to pull closer the matched phrase-object pairs and push away unmatched ones for aggregating more precise object contexts from more phrase-relevant object tokens. Extensive experiments on the PNG benchmark show our method achieves new state-of-the-art performance with large margins.
title Enriching Phrases with Coupled Pixel and Object Contexts for Panoptic Narrative Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.01091