OpenVIS: Open-vocabulary Video Instance Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Pinxue, Huang, Tony, He, Peiyang, Liu, Xuefeng, Xiao, Tianjun, Chen, Zhaoyu, Zhang, Wenqiang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910568378531840
author Guo, Pinxue
Huang, Tony
He, Peiyang
Liu, Xuefeng
Xiao, Tianjun
Chen, Zhaoyu
Zhang, Wenqiang
author_facet Guo, Pinxue
Huang, Tony
He, Peiyang
Liu, Xuefeng
Xiao, Tianjun
Chen, Zhaoyu
Zhang, Wenqiang
contents Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose InstFormer, a carefully designed framework for the OpenVIS task that achieves powerful open-vocabulary capabilities through lightweight fine-tuning with limited-category data. InstFormer begins with the open-world mask proposal network, encouraged to propose all potential instance class-agnostic masks by the contrastive instance margin loss. Next, we introduce InstCLIP, adapted from pre-trained CLIP with Instance Guidance Attention, which encodes open-vocabulary instance tokens efficiently. These instance tokens not only enable open-vocabulary classification but also offer strong universal tracking capabilities. Furthermore, to prevent the tracking module from being constrained by the training data with limited categories, we propose the universal rollout association, which transforms the tracking problem into predicting the next frame's instance tracking token. The experimental results demonstrate the proposed InstFormer achieve state-of-the-art capabilities on a comprehensive OpenVIS evaluation benchmark, while also achieves competitive performance in fully supervised VIS task.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16835
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle OpenVIS: Open-vocabulary Video Instance Segmentation
Guo, Pinxue
Huang, Tony
He, Peiyang
Liu, Xuefeng
Xiao, Tianjun
Chen, Zhaoyu
Zhang, Wenqiang
Computer Vision and Pattern Recognition
Open-vocabulary Video Instance Segmentation (OpenVIS) can simultaneously detect, segment, and track arbitrary object categories in a video, without being constrained to categories seen during training. In this work, we propose InstFormer, a carefully designed framework for the OpenVIS task that achieves powerful open-vocabulary capabilities through lightweight fine-tuning with limited-category data. InstFormer begins with the open-world mask proposal network, encouraged to propose all potential instance class-agnostic masks by the contrastive instance margin loss. Next, we introduce InstCLIP, adapted from pre-trained CLIP with Instance Guidance Attention, which encodes open-vocabulary instance tokens efficiently. These instance tokens not only enable open-vocabulary classification but also offer strong universal tracking capabilities. Furthermore, to prevent the tracking module from being constrained by the training data with limited categories, we propose the universal rollout association, which transforms the tracking problem into predicting the next frame's instance tracking token. The experimental results demonstrate the proposed InstFormer achieve state-of-the-art capabilities on a comprehensive OpenVIS evaluation benchmark, while also achieves competitive performance in fully supervised VIS task.
title OpenVIS: Open-vocabulary Video Instance Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2305.16835