Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929518201012224 |
|---|---|
| author | Sidhu, Mankeerat Chopra, Hetarth Blume, Ansel Kim, Jeonghwan Reddy, Revanth Gangi Ji, Heng |
| author_facet | Sidhu, Mankeerat Chopra, Hetarth Blume, Ansel Kim, Jeonghwan Reddy, Revanth Gangi Ji, Heng |
| contents | In this paper, we introduce SearchDet, a training-free long-tail object detection framework that significantly enhances open-vocabulary object detection performance. SearchDet retrieves a set of positive and negative images of an object to ground, embeds these images, and computes an input image-weighted query which is used to detect the desired concept in the image. Our proposed method is simple and training-free, yet achieves over 48.7% mAP improvement on ODinW and 59.1% mAP improvement on LVIS compared to state-of-the-art models such as GroundingDINO. We further show that our approach of basing object detection on a set of Web-retrieved exemplars is stable with respect to variations in the exemplars, suggesting a path towards eliminating costly data annotation and training procedures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_18733 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval Sidhu, Mankeerat Chopra, Hetarth Blume, Ansel Kim, Jeonghwan Reddy, Revanth Gangi Ji, Heng Computer Vision and Pattern Recognition In this paper, we introduce SearchDet, a training-free long-tail object detection framework that significantly enhances open-vocabulary object detection performance. SearchDet retrieves a set of positive and negative images of an object to ground, embeds these images, and computes an input image-weighted query which is used to detect the desired concept in the image. Our proposed method is simple and training-free, yet achieves over 48.7% mAP improvement on ODinW and 59.1% mAP improvement on LVIS compared to state-of-the-art models such as GroundingDINO. We further show that our approach of basing object detection on a set of Web-retrieved exemplars is stable with respect to variations in the exemplars, suggesting a path towards eliminating costly data annotation and training procedures. |
| title | Search and Detect: Training-Free Long Tail Object Detection via Web-Image Retrieval |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.18733 |