Guided Slot Attention for Unsupervised Video Object Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Minhyeok, Cho, Suhwan, Lee, Dogyoon, Park, Chaewon, Lee, Jungho, Lee, Sangyoun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909155121430528
author Lee, Minhyeok
Cho, Suhwan
Lee, Dogyoon
Park, Chaewon
Lee, Jungho
Lee, Sangyoun
author_facet Lee, Minhyeok
Cho, Suhwan
Lee, Dogyoon
Park, Chaewon
Lee, Jungho
Lee, Sangyoun
contents Unsupervised video object segmentation aims to segment the most prominent object in a video sequence. However, the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue, we propose a guided slot attention network to reinforce spatial structural information and obtain better foreground--background separation. The foreground and background slots, which are initialized with query guidance, are iteratively refined based on interactions with template information. Furthermore, to improve slot--template interaction and effectively fuse global and local features in the target and reference frames, K-nearest neighbors filtering and a feature aggregation transformer are introduced. The proposed model achieves state-of-the-art performance on two popular datasets. Additionally, we demonstrate the robustness of the proposed model in challenging scenes through various comparative experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2303_08314
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Guided Slot Attention for Unsupervised Video Object Segmentation
Lee, Minhyeok
Cho, Suhwan
Lee, Dogyoon
Park, Chaewon
Lee, Jungho
Lee, Sangyoun
Computer Vision and Pattern Recognition
Unsupervised video object segmentation aims to segment the most prominent object in a video sequence. However, the existence of complex backgrounds and multiple foreground objects make this task challenging. To address this issue, we propose a guided slot attention network to reinforce spatial structural information and obtain better foreground--background separation. The foreground and background slots, which are initialized with query guidance, are iteratively refined based on interactions with template information. Furthermore, to improve slot--template interaction and effectively fuse global and local features in the target and reference frames, K-nearest neighbors filtering and a feature aggregation transformer are introduced. The proposed model achieves state-of-the-art performance on two popular datasets. Additionally, we demonstrate the robustness of the proposed model in challenging scenes through various comparative experiments.
title Guided Slot Attention for Unsupervised Video Object Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.08314