Taming Self-Training for Open-Vocabulary Object Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Shiyu, Schulter, Samuel, Zhao, Long, Zhang, Zhixing, G, Vijay Kumar B., Suh, Yumin, Chandraker, Manmohan, Metaxas, Dimitris N.
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910408644755456
author Zhao, Shiyu
Schulter, Samuel
Zhao, Long
Zhang, Zhixing
G, Vijay Kumar B.
Suh, Yumin
Chandraker, Manmohan
Metaxas, Dimitris N.
author_facet Zhao, Shiyu
Schulter, Samuel
Zhao, Long
Zhang, Zhixing
G, Vijay Kumar B.
Suh, Yumin
Chandraker, Manmohan
Metaxas, Dimitris N.
contents Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pretrained vision and language models (VLMs). However, teacher-student self-training, a powerful and widely used paradigm to leverage PLs, is rarely explored for OVD. This work identifies two challenges of using self-training in OVD: noisy PLs from VLMs and frequent distribution changes of PLs. To address these challenges, we propose SAS-Det that tames self-training for OVD from two key perspectives. First, we present a split-and-fusion (SAF) head that splits a standard detection into an open-branch and a closed-branch. This design can reduce noisy supervision from pseudo boxes. Moreover, the two branches learn complementary knowledge from different training data, significantly enhancing performance when fused together. Second, in our view, unlike in closed-set tasks, the PL distributions in OVD are solely determined by the teacher model. We introduce a periodic update strategy to decrease the number of updates to the teacher, thereby decreasing the frequency of changes in PL distributions, which stabilizes the training process. Extensive experiments demonstrate SAS-Det is both efficient and effective. SAS-Det outperforms recent models of the same scale by a clear margin and achieves 37.4 AP50 and 29.1 APr on novel categories of the COCO and LVIS benchmarks, respectively. Code is available at \url{https://github.com/xiaofeng94/SAS-Det}.
format Preprint
id arxiv_https___arxiv_org_abs_2308_06412
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Taming Self-Training for Open-Vocabulary Object Detection
Zhao, Shiyu
Schulter, Samuel
Zhao, Long
Zhang, Zhixing
G, Vijay Kumar B.
Suh, Yumin
Chandraker, Manmohan
Metaxas, Dimitris N.
Computer Vision and Pattern Recognition
Recent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pretrained vision and language models (VLMs). However, teacher-student self-training, a powerful and widely used paradigm to leverage PLs, is rarely explored for OVD. This work identifies two challenges of using self-training in OVD: noisy PLs from VLMs and frequent distribution changes of PLs. To address these challenges, we propose SAS-Det that tames self-training for OVD from two key perspectives. First, we present a split-and-fusion (SAF) head that splits a standard detection into an open-branch and a closed-branch. This design can reduce noisy supervision from pseudo boxes. Moreover, the two branches learn complementary knowledge from different training data, significantly enhancing performance when fused together. Second, in our view, unlike in closed-set tasks, the PL distributions in OVD are solely determined by the teacher model. We introduce a periodic update strategy to decrease the number of updates to the teacher, thereby decreasing the frequency of changes in PL distributions, which stabilizes the training process. Extensive experiments demonstrate SAS-Det is both efficient and effective. SAS-Det outperforms recent models of the same scale by a clear margin and achieves 37.4 AP50 and 29.1 APr on novel categories of the COCO and LVIS benchmarks, respectively. Code is available at \url{https://github.com/xiaofeng94/SAS-Det}.
title Taming Self-Training for Open-Vocabulary Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.06412