A Unified Structure for Efficient RGB and RGB-D Salient Object Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Peng, Peng, Li, Yong-Jie
Natura: Preprint
Pubblicazione: 2020
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929565161488384
author Peng, Peng
Li, Yong-Jie
author_facet Peng, Peng
Li, Yong-Jie
contents Salient object detection (SOD) has been well studied in recent years, especially using deep neural networks. However, SOD with RGB and RGB-D images is usually treated as two different tasks with different network structures that need to be designed specifically. In this paper, we proposed a unified and efficient structure with a cross-attention context extraction (CRACE) module to address both tasks of SOD efficiently. The proposed CRACE module receives and appropriately fuses two (for RGB SOD) or three (for RGB-D SOD) inputs. The simple unified feature pyramid network (FPN)-like structure with CRACE modules conveys and refines the results under the multi-level supervisions of saliency and boundaries. The proposed structure is simple yet effective; the rich context information of RGB and depth can be appropriately extracted and fused by the proposed structure efficiently. Experimental results show that our method outperforms other state-of-the-art methods in both RGB and RGB-D SOD tasks on various datasets and in terms of most metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2012_00437
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle A Unified Structure for Efficient RGB and RGB-D Salient Object Detection
Peng, Peng
Li, Yong-Jie
Computer Vision and Pattern Recognition
Salient object detection (SOD) has been well studied in recent years, especially using deep neural networks. However, SOD with RGB and RGB-D images is usually treated as two different tasks with different network structures that need to be designed specifically. In this paper, we proposed a unified and efficient structure with a cross-attention context extraction (CRACE) module to address both tasks of SOD efficiently. The proposed CRACE module receives and appropriately fuses two (for RGB SOD) or three (for RGB-D SOD) inputs. The simple unified feature pyramid network (FPN)-like structure with CRACE modules conveys and refines the results under the multi-level supervisions of saliency and boundaries. The proposed structure is simple yet effective; the rich context information of RGB and depth can be appropriately extracted and fused by the proposed structure efficiently. Experimental results show that our method outperforms other state-of-the-art methods in both RGB and RGB-D SOD tasks on various datasets and in terms of most metrics.
title A Unified Structure for Efficient RGB and RGB-D Salient Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2012.00437