Phrase-Instance Alignment for Generalized Referring Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nguyen, E-Ro, Le, Hieu, Samaras, Dimitris, Ryoo, Michael S.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918407196114944
author Nguyen, E-Ro
Le, Hieu
Samaras, Dimitris
Ryoo, Michael S.
author_facet Nguyen, E-Ro
Le, Hieu
Samaras, Dimitris
Ryoo, Michael S.
contents Generalized Referring expressions can describe one object, several related objects, or none at all. Existing generalized referring segmentation (GRES) models treat all cases alike, predicting a single binary mask and ignoring how linguistic phrases correspond to distinct visual instances. To this end, we reformulate GRES as an instance-level reasoning problem, where the model first predicts multiple instance-aware object queries conditioned on the referring expression, then aligns each with its most relevant phrase. This alignment is enforced by a Phrase-Object Alignment (POA) loss that builds fine-grained correspondence between linguistic phrases and visual instances. Given these aligned object instance queries and their learned relevance scores, the final segmentation and the no-target case are both inferred through a unified relevance-weighted aggregation mechanism. This instance-aware formulation enables explicit phrase-instance grounding, interpretable reasoning, and robust handling of complex or null expressions. Extensive experiments on the gRefCOCO and Ref-ZOM benchmarks demonstrate that our method significantly advances state-of-the-art performance by 3.22% cIoU and 12.25% N-acc.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15087
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Phrase-Instance Alignment for Generalized Referring Segmentation
Nguyen, E-Ro
Le, Hieu
Samaras, Dimitris
Ryoo, Michael S.
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Generalized Referring expressions can describe one object, several related objects, or none at all. Existing generalized referring segmentation (GRES) models treat all cases alike, predicting a single binary mask and ignoring how linguistic phrases correspond to distinct visual instances. To this end, we reformulate GRES as an instance-level reasoning problem, where the model first predicts multiple instance-aware object queries conditioned on the referring expression, then aligns each with its most relevant phrase. This alignment is enforced by a Phrase-Object Alignment (POA) loss that builds fine-grained correspondence between linguistic phrases and visual instances. Given these aligned object instance queries and their learned relevance scores, the final segmentation and the no-target case are both inferred through a unified relevance-weighted aggregation mechanism. This instance-aware formulation enables explicit phrase-instance grounding, interpretable reasoning, and robust handling of complex or null expressions. Extensive experiments on the gRefCOCO and Ref-ZOM benchmarks demonstrate that our method significantly advances state-of-the-art performance by 3.22% cIoU and 12.25% N-acc.
title Phrase-Instance Alignment for Generalized Referring Segmentation
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.15087