Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shilong, Zeng, Zhaoyang, Ren, Tianhe, Li, Feng, Zhang, Hao, Yang, Jie, Jiang, Qing, Li, Chunyuan, Yang, Jianwei, Su, Hang, Zhu, Jun, Zhang, Lei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917727200870400
author Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Hao
Yang, Jie
Jiang, Qing
Li, Chunyuan
Yang, Jianwei
Su, Hang
Zhu, Jun
Zhang, Lei
author_facet Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Hao
Yang, Jie
Jiang, Qing
Li, Chunyuan
Yang, Jianwei
Su, Hang
Zhu, Jun
Zhang, Lei
contents In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection is introducing language to a closed-set detector for open-set concept generalization. To effectively fuse language and vision modalities, we conceptually divide a closed-set detector into three phases and propose a tight fusion solution, which includes a feature enhancer, a language-guided query selection, and a cross-modality decoder for cross-modality fusion. While previous works mainly evaluate open-set object detection on novel categories, we propose to also perform evaluations on referring expression comprehension for objects specified with attributes. Grounding DINO performs remarkably well on all three settings, including benchmarks on COCO, LVIS, ODinW, and RefCOCO/+/g. Grounding DINO achieves a $52.5$ AP on the COCO detection zero-shot transfer benchmark, i.e., without any training data from COCO. It sets a new record on the ODinW zero-shot benchmark with a mean $26.1$ AP. Code will be available at \url{https://github.com/IDEA-Research/GroundingDINO}.
format Preprint
id arxiv_https___arxiv_org_abs_2303_05499
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Liu, Shilong
Zeng, Zhaoyang
Ren, Tianhe
Li, Feng
Zhang, Hao
Yang, Jie
Jiang, Qing
Li, Chunyuan
Yang, Jianwei
Su, Hang
Zhu, Jun
Zhang, Lei
Computer Vision and Pattern Recognition
In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection is introducing language to a closed-set detector for open-set concept generalization. To effectively fuse language and vision modalities, we conceptually divide a closed-set detector into three phases and propose a tight fusion solution, which includes a feature enhancer, a language-guided query selection, and a cross-modality decoder for cross-modality fusion. While previous works mainly evaluate open-set object detection on novel categories, we propose to also perform evaluations on referring expression comprehension for objects specified with attributes. Grounding DINO performs remarkably well on all three settings, including benchmarks on COCO, LVIS, ODinW, and RefCOCO/+/g. Grounding DINO achieves a $52.5$ AP on the COCO detection zero-shot transfer benchmark, i.e., without any training data from COCO. It sets a new record on the ODinW zero-shot benchmark with a mean $26.1$ AP. Code will be available at \url{https://github.com/IDEA-Research/GroundingDINO}.
title Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2303.05499