PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Xu, Yang, Kailun, Lin, Jiacheng, Yuan, Jin, Li, Zhiyong, Li, Shutao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913569863368704
author Zhang, Xu
Yang, Kailun
Lin, Jiacheng
Yuan, Jin
Li, Zhiyong
Li, Shutao
author_facet Zhang, Xu
Yang, Kailun
Lin, Jiacheng
Yuan, Jin
Li, Zhiyong
Li, Shutao
contents Integration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interaction efficiency. However, existing studies primarily encode the position or pixel regions of prompts without considering the contextual areas around them, resulting in insufficient prompt feedback, which is not conducive to performance acceleration. To tackle this problem, this paper proposes a simple yet effective Probabilistic Visual Prompt Unified Transformer (PVPUFormer) for interactive image segmentation, which allows users to flexibly input diverse visual prompts with the probabilistic prompt encoding and feature post-processing to excavate sufficient and robust prompt features for performance boosting. Specifically, we first propose a Probabilistic Prompt-unified Encoder (PPuE) to generate a unified one-dimensional vector by exploring both prompt and non-prompt contextual information, offering richer feedback cues to accelerate performance improvement. On this basis, we further present a Prompt-to-Pixel Contrastive (P$^2$C) loss to accurately align both prompt and pixel features, bridging the representation gap between them to offer consistent feature representations for mask prediction. Moreover, our approach designs a Dual-cross Merging Attention (DMA) module to implement bidirectional feature interaction between image and prompt features, generating notable features for performance improvement. A comprehensive variety of experiments on several challenging datasets demonstrates that the proposed components achieve consistent improvements, yielding state-of-the-art interactive segmentation performance. Our code is available at https://github.com/XuZhang1211/PVPUFormer.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06656
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
Zhang, Xu
Yang, Kailun
Lin, Jiacheng
Yuan, Jin
Li, Zhiyong
Li, Shutao
Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
Integration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interaction efficiency. However, existing studies primarily encode the position or pixel regions of prompts without considering the contextual areas around them, resulting in insufficient prompt feedback, which is not conducive to performance acceleration. To tackle this problem, this paper proposes a simple yet effective Probabilistic Visual Prompt Unified Transformer (PVPUFormer) for interactive image segmentation, which allows users to flexibly input diverse visual prompts with the probabilistic prompt encoding and feature post-processing to excavate sufficient and robust prompt features for performance boosting. Specifically, we first propose a Probabilistic Prompt-unified Encoder (PPuE) to generate a unified one-dimensional vector by exploring both prompt and non-prompt contextual information, offering richer feedback cues to accelerate performance improvement. On this basis, we further present a Prompt-to-Pixel Contrastive (P$^2$C) loss to accurately align both prompt and pixel features, bridging the representation gap between them to offer consistent feature representations for mask prediction. Moreover, our approach designs a Dual-cross Merging Attention (DMA) module to implement bidirectional feature interaction between image and prompt features, generating notable features for performance improvement. A comprehensive variety of experiments on several challenging datasets demonstrates that the proposed components achieve consistent improvements, yielding state-of-the-art interactive segmentation performance. Our code is available at https://github.com/XuZhang1211/PVPUFormer.
title PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
topic Computer Vision and Pattern Recognition
Robotics
Image and Video Processing
url https://arxiv.org/abs/2306.06656