InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Feng, Chengjian, Zhong, Yujie, Jie, Zequn, Xie, Weidi, Ma, Lin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914743199989760
author Feng, Chengjian
Zhong, Yujie
Jie, Zequn
Xie, Weidi
Ma, Lin
author_facet Feng, Chengjian
Zhong, Yujie
Jie, Zequn
Xie, Weidi
Ma, Lin
contents In this paper, we present a novel paradigm to enhance the ability of object detector, e.g., expanding categories or improving detection performance, by training on synthetic dataset generated from diffusion models. Specifically, we integrate an instance-level grounding head into a pre-trained, generative diffusion model, to augment it with the ability of localising instances in the generated images. The grounding head is trained to align the text embedding of category names with the regional visual feature of the diffusion model, using supervision from an off-the-shelf object detector, and a novel self-training scheme on (novel) categories not covered by the detector. We conduct thorough experiments to show that, this enhanced version of diffusion model, termed as InstaGen, can serve as a data synthesizer, to enhance object detectors by training on its generated samples, demonstrating superior performance over existing state-of-the-art methods in open-vocabulary (+4.5 AP) and data-sparse (+1.2 to 5.2 AP) scenarios. Project page with code: https://fcjian.github.io/InstaGen.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05937
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
Feng, Chengjian
Zhong, Yujie
Jie, Zequn
Xie, Weidi
Ma, Lin
Computer Vision and Pattern Recognition
In this paper, we present a novel paradigm to enhance the ability of object detector, e.g., expanding categories or improving detection performance, by training on synthetic dataset generated from diffusion models. Specifically, we integrate an instance-level grounding head into a pre-trained, generative diffusion model, to augment it with the ability of localising instances in the generated images. The grounding head is trained to align the text embedding of category names with the regional visual feature of the diffusion model, using supervision from an off-the-shelf object detector, and a novel self-training scheme on (novel) categories not covered by the detector. We conduct thorough experiments to show that, this enhanced version of diffusion model, termed as InstaGen, can serve as a data synthesizer, to enhance object detectors by training on its generated samples, demonstrating superior performance over existing state-of-the-art methods in open-vocabulary (+4.5 AP) and data-sparse (+1.2 to 5.2 AP) scenarios. Project page with code: https://fcjian.github.io/InstaGen.
title InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.05937