RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866916266022797312 |
|---|---|
| author | Chen, Fangyi Zhang, Han Yang, Zhantao Chen, Hao Hu, Kai Savvides, Marios |
| author_facet | Chen, Fangyi Zhang, Han Yang, Zhantao Chen, Hao Hu, Kai Savvides, Marios |
| contents | Open-vocabulary object detection (OVD) requires solid modeling of the region-semantic relationship, which could be learned from massive region-text pairs. However, such data is limited in practice due to significant annotation costs. In this work, we propose RTGen to generate scalable open-vocabulary region-text pairs and demonstrate its capability to boost the performance of open-vocabulary object detection. RTGen includes both text-to-region and region-to-text generation processes on scalable image-caption data. The text-to-region generation is powered by image inpainting, directed by our proposed scene-aware inpainting guider for overall layout harmony. For region-to-text generation, we perform multiple region-level image captioning with various prompts and select the best matching text according to CLIP similarity. To facilitate detection training on region-text pairs, we also introduce a localization-aware region-text contrastive loss that learns object proposals tailored with different localization qualities. Extensive experiments demonstrate that our RTGen can serve as a scalable, semantically rich, and effective source for open-vocabulary object detection and continue to improve the model performance when more data is utilized, delivering superior performance compared to the existing state-of-the-art methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_19854 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection Chen, Fangyi Zhang, Han Yang, Zhantao Chen, Hao Hu, Kai Savvides, Marios Computer Vision and Pattern Recognition Open-vocabulary object detection (OVD) requires solid modeling of the region-semantic relationship, which could be learned from massive region-text pairs. However, such data is limited in practice due to significant annotation costs. In this work, we propose RTGen to generate scalable open-vocabulary region-text pairs and demonstrate its capability to boost the performance of open-vocabulary object detection. RTGen includes both text-to-region and region-to-text generation processes on scalable image-caption data. The text-to-region generation is powered by image inpainting, directed by our proposed scene-aware inpainting guider for overall layout harmony. For region-to-text generation, we perform multiple region-level image captioning with various prompts and select the best matching text according to CLIP similarity. To facilitate detection training on region-text pairs, we also introduce a localization-aware region-text contrastive loss that learns object proposals tailored with different localization qualities. Extensive experiments demonstrate that our RTGen can serve as a scalable, semantically rich, and effective source for open-vocabulary object detection and continue to improve the model performance when more data is utilized, delivering superior performance compared to the existing state-of-the-art methods. |
| title | RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.19854 |