VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Luo, Ziyang, Liu, Nian, Zhao, Wangbo, Yang, Xuguang, Zhang, Dingwen, Fan, Deng-Ping, Khan, Fahad, Han, Junwei
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916200556003328
author Luo, Ziyang
Liu, Nian
Zhao, Wangbo
Yang, Xuguang
Zhang, Dingwen
Fan, Deng-Ping
Khan, Fahad
Han, Junwei
author_facet Luo, Ziyang
Liu, Nian
Zhao, Wangbo
Yang, Xuguang
Zhang, Dingwen
Fan, Deng-Ping
Khan, Fahad
Han, Junwei
contents Salient object detection (SOD) and camouflaged object detection (COD) are related yet distinct binary mapping tasks. These tasks involve multiple modalities, sharing commonalities and unique cues. Existing research often employs intricate task-specific specialist models, potentially leading to redundancy and suboptimal results. We introduce VSCode, a generalist model with novel 2D prompt learning, to jointly address four SOD tasks and three COD tasks. We utilize VST as the foundation model and introduce 2D prompts within the encoder-decoder architecture to learn domain and task-specific knowledge on two separate dimensions. A prompt discrimination loss helps disentangle peculiarities to benefit model optimization. VSCode outperforms state-of-the-art methods across six tasks on 26 datasets and exhibits zero-shot generalization to unseen tasks by combining 2D prompts, such as RGB-D COD. Source code has been available at https://github.com/Sssssuperior/VSCode.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15011
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning
Luo, Ziyang
Liu, Nian
Zhao, Wangbo
Yang, Xuguang
Zhang, Dingwen
Fan, Deng-Ping
Khan, Fahad
Han, Junwei
Computer Vision and Pattern Recognition
Salient object detection (SOD) and camouflaged object detection (COD) are related yet distinct binary mapping tasks. These tasks involve multiple modalities, sharing commonalities and unique cues. Existing research often employs intricate task-specific specialist models, potentially leading to redundancy and suboptimal results. We introduce VSCode, a generalist model with novel 2D prompt learning, to jointly address four SOD tasks and three COD tasks. We utilize VST as the foundation model and introduce 2D prompts within the encoder-decoder architecture to learn domain and task-specific knowledge on two separate dimensions. A prompt discrimination loss helps disentangle peculiarities to benefit model optimization. VSCode outperforms state-of-the-art methods across six tasks on 26 datasets and exhibits zero-shot generalization to unseen tasks by combining 2D prompts, such as RGB-D COD. Source code has been available at https://github.com/Sssssuperior/VSCode.
title VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15011