An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Xiangyu, Chen, Yicheng, Xu, Shilin, Li, Xiangtai, Wang, Xinjiang, Li, Yining, Huang, Haian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910287718776832
author Zhao, Xiangyu
Chen, Yicheng
Xu, Shilin
Li, Xiangtai
Wang, Xinjiang
Li, Yining
Huang, Haian
author_facet Zhao, Xiangyu
Chen, Yicheng
Xu, Shilin
Li, Xiangtai
Wang, Xinjiang
Li, Yining
Huang, Haian
contents Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to its widespread adoption as a mainstream architecture for various downstream applications. However, despite its significance, the original Grounding-DINO model lacks comprehensive public technical details due to the unavailability of its training code. To bridge this gap, we present MM-Grounding-DINO, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox. It adopts abundant vision datasets for pre-training and various detection and grounding datasets for fine-tuning. We give a comprehensive analysis of each reported result and detailed settings for reproduction. The extensive experiments on the benchmarks mentioned demonstrate that our MM-Grounding-DINO-Tiny outperforms the Grounding-DINO-Tiny baseline. We release all our models to the research community. Codes and trained models are released at https://github.com/open-mmlab/mmdetection/tree/main/configs/mm_grounding_dino.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02361
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
Zhao, Xiangyu
Chen, Yicheng
Xu, Shilin
Li, Xiangtai
Wang, Xinjiang
Li, Yining
Huang, Haian
Computer Vision and Pattern Recognition
Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to its widespread adoption as a mainstream architecture for various downstream applications. However, despite its significance, the original Grounding-DINO model lacks comprehensive public technical details due to the unavailability of its training code. To bridge this gap, we present MM-Grounding-DINO, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox. It adopts abundant vision datasets for pre-training and various detection and grounding datasets for fine-tuning. We give a comprehensive analysis of each reported result and detailed settings for reproduction. The extensive experiments on the benchmarks mentioned demonstrate that our MM-Grounding-DINO-Tiny outperforms the Grounding-DINO-Tiny baseline. We release all our models to the research community. Codes and trained models are released at https://github.com/open-mmlab/mmdetection/tree/main/configs/mm_grounding_dino.
title An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.02361