UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zhenghao, Zhang, Shengfan, Wei, Zhichao, Dai, Zuozhuo, Zhu, Siyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912469638709248
author Zhang, Zhenghao
Zhang, Shengfan
Wei, Zhichao
Dai, Zuozhuo
Zhu, Siyu
author_facet Zhang, Zhenghao
Zhang, Shengfan
Wei, Zhichao
Dai, Zuozhuo
Zhu, Siyu
contents The current state-of-the-art methods for unsupervised video object segmentation (UVOS) require extensive training on video datasets with mask annotations, limiting their effectiveness in handling challenging scenarios. However, the Segment Anything Model (SAM) introduces a new prompt-driven paradigm for image segmentation, offering new possibilities. In this study, we investigate SAM's potential for UVOS through different prompt strategies. We then propose UVOSAM, a mask-free paradigm for UVOS that utilizes the STD-Net tracker. STD-Net incorporates a spatial-temporal decoupled deformable attention mechanism to establish an effective correlation between intra- and inter-frame features, remarkably enhancing the quality of box prompts in complex video scenes. Extensive experiments on the DAVIS2017-unsupervised and YoutubeVIS19\&21 datasets demonstrate the superior performance of UVOSAM without mask supervision compared to existing mask-supervised methods, as well as its ability to generalize to weakly-annotated video datasets. Code can be found at https://github.com/alibaba/UVOSAM.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12659
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
Zhang, Zhenghao
Zhang, Shengfan
Wei, Zhichao
Dai, Zuozhuo
Zhu, Siyu
Computer Vision and Pattern Recognition
The current state-of-the-art methods for unsupervised video object segmentation (UVOS) require extensive training on video datasets with mask annotations, limiting their effectiveness in handling challenging scenarios. However, the Segment Anything Model (SAM) introduces a new prompt-driven paradigm for image segmentation, offering new possibilities. In this study, we investigate SAM's potential for UVOS through different prompt strategies. We then propose UVOSAM, a mask-free paradigm for UVOS that utilizes the STD-Net tracker. STD-Net incorporates a spatial-temporal decoupled deformable attention mechanism to establish an effective correlation between intra- and inter-frame features, remarkably enhancing the quality of box prompts in complex video scenes. Extensive experiments on the DAVIS2017-unsupervised and YoutubeVIS19\&21 datasets demonstrate the superior performance of UVOSAM without mask supervision compared to existing mask-supervised methods, as well as its ability to generalize to weakly-annotated video datasets. Code can be found at https://github.com/alibaba/UVOSAM.
title UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.12659