Full-Duplex Strategy for Video Object Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ji, Ge-Peng, Fan, Deng-Ping, Fu, Keren, Wu, Zhe, Shen, Jianbing, Shao, Ling
Format: Preprint
Veröffentlicht: 2021
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916130228011008
author Ji, Ge-Peng
Fan, Deng-Ping
Fu, Keren
Wu, Zhe
Shen, Jianbing
Shao, Ling
author_facet Ji, Ge-Peng
Fan, Deng-Ping
Fu, Keren
Wu, Zhe
Shen, Jianbing
Shao, Ling
contents Previous video object segmentation approaches mainly focus on using simplex solutions between appearance and motion, limiting feature collaboration efficiency among and across these two cues. In this work, we study a novel and efficient full-duplex strategy network (FSNet) to address this issue, by considering a better mutual restraint scheme between motion and appearance in exploiting the cross-modal features from the fusion and decoding stage. Specifically, we introduce the relational cross-attention module (RCAM) to achieve bidirectional message propagation across embedding sub-spaces. To improve the model's robustness and update the inconsistent features from the spatial-temporal embeddings, we adopt the bidirectional purification module (BPM) after the RCAM. Extensive experiments on five popular benchmarks show that our FSNet is robust to various challenging scenarios (e.g., motion blur, occlusion) and achieves favourable performance against existing cutting-edges both in the video object segmentation and video salient object detection tasks. The project is publicly available at: https://dpfan.net/FSNet.
format Preprint
id arxiv_https___arxiv_org_abs_2108_03151
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Full-Duplex Strategy for Video Object Segmentation
Ji, Ge-Peng
Fan, Deng-Ping
Fu, Keren
Wu, Zhe
Shen, Jianbing
Shao, Ling
Computer Vision and Pattern Recognition
Previous video object segmentation approaches mainly focus on using simplex solutions between appearance and motion, limiting feature collaboration efficiency among and across these two cues. In this work, we study a novel and efficient full-duplex strategy network (FSNet) to address this issue, by considering a better mutual restraint scheme between motion and appearance in exploiting the cross-modal features from the fusion and decoding stage. Specifically, we introduce the relational cross-attention module (RCAM) to achieve bidirectional message propagation across embedding sub-spaces. To improve the model's robustness and update the inconsistent features from the spatial-temporal embeddings, we adopt the bidirectional purification module (BPM) after the RCAM. Extensive experiments on five popular benchmarks show that our FSNet is robust to various challenging scenarios (e.g., motion blur, occlusion) and achieves favourable performance against existing cutting-edges both in the video object segmentation and video salient object detection tasks. The project is publicly available at: https://dpfan.net/FSNet.
title Full-Duplex Strategy for Video Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2108.03151