LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913672909029376 |
|---|---|
| author | Fu, Shenghao Yang, Qize Mo, Qijie Yan, Junkai Wei, Xihan Meng, Jingke Xie, Xiaohua Zheng, Wei-Shi |
| author_facet | Fu, Shenghao Yang, Qize Mo, Qijie Yan, Junkai Wei, Xihan Meng, Jingke Xie, Xiaohua Zheng, Wei-Shi |
| contents | Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https://github.com/iSEE-Laboratory/LLMDet. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_18954 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models Fu, Shenghao Yang, Qize Mo, Qijie Yan, Junkai Wei, Xihan Meng, Jingke Xie, Xiaohua Zheng, Wei-Shi Computer Vision and Pattern Recognition Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achieve the goal, we first collect a dataset, GroundingCap-1M, wherein each image is accompanied by associated grounding labels and an image-level detailed caption. With this dataset, we finetune an open-vocabulary detector with training objectives including a standard grounding loss and a caption generation loss. We take advantage of a large language model to generate both region-level short captions for each region of interest and image-level long captions for the whole image. Under the supervision of the large language model, the resulting detector, LLMDet, outperforms the baseline by a clear margin, enjoying superior open-vocabulary ability. Further, we show that the improved LLMDet can in turn build a stronger large multi-modal model, achieving mutual benefits. The code, model, and dataset is available at https://github.com/iSEE-Laboratory/LLMDet. |
| title | LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.18954 |