Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Miao, Deshui, Gu, Yameng, Li, Xin, He, Zhenyu, Wang, Yaowei, Yang, Ming-Hsuan
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914928845127680
author Miao, Deshui
Gu, Yameng
Li, Xin
He, Zhenyu
Wang, Yaowei
Yang, Ming-Hsuan
author_facet Miao, Deshui
Gu, Yameng
Li, Xin
He, Zhenyu
Wang, Yaowei
Yang, Ming-Hsuan
contents Video object segmentation (VOS) is a crucial task in computer vision, but current VOS methods struggle with complex scenes and prolonged object motions. To address these challenges, the MOSE dataset aims to enhance object recognition and differentiation in complex environments, while the LVOS dataset focuses on segmenting objects exhibiting long-term, intricate movements. This report introduces a discriminative spatial-temporal VOS model that utilizes discriminative object features as query representations. The semantic understanding of spatial-semantic modules enables it to recognize object parts, while salient features highlight more distinctive object characteristics. Our model, trained on extensive VOS datasets, achieved first place (\textbf{80.90\%} $\mathcal{J \& F}$) on the test set of the 6th LSVOS challenge in the VOS Track, demonstrating its effectiveness in tackling the aforementioned challenges. The code will be available at \href{https://github.com/yahooo-m/VOS-Solution}{code}.
format Preprint
id arxiv_https___arxiv_org_abs_2408_16431
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS
Miao, Deshui
Gu, Yameng
Li, Xin
He, Zhenyu
Wang, Yaowei
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
Video object segmentation (VOS) is a crucial task in computer vision, but current VOS methods struggle with complex scenes and prolonged object motions. To address these challenges, the MOSE dataset aims to enhance object recognition and differentiation in complex environments, while the LVOS dataset focuses on segmenting objects exhibiting long-term, intricate movements. This report introduces a discriminative spatial-temporal VOS model that utilizes discriminative object features as query representations. The semantic understanding of spatial-semantic modules enables it to recognize object parts, while salient features highlight more distinctive object characteristics. Our model, trained on extensive VOS datasets, achieved first place (\textbf{80.90\%} $\mathcal{J \& F}$) on the test set of the 6th LSVOS challenge in the VOS Track, demonstrating its effectiveness in tackling the aforementioned challenges. The code will be available at \href{https://github.com/yahooo-m/VOS-Solution}{code}.
title Discriminative Spatial-Semantic VOS Solution: 1st Place Solution for 6th LSVOS
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.16431