CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhichao, Hu, Huazhang, Ma, Yidong, Liu, Gang, Chen, Yibo, Tang, Xu, Hu, Yao, Xu, Yongchao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911200560807936
author Sun, Zhichao
Hu, Huazhang
Ma, Yidong
Liu, Gang
Chen, Yibo
Tang, Xu
Hu, Yao
Xu, Yongchao
author_facet Sun, Zhichao
Hu, Huazhang
Ma, Yidong
Liu, Gang
Chen, Yibo
Tang, Xu
Hu, Yao
Xu, Yongchao
contents With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive insufficient learning signals, and hard negative gradient dilution, where discriminative gradients are overwhelmed by numerous easy negatives. To address these challenges, we propose CQ-DINO, a category query-based object detection framework that reformulates classification as a contrastive task between object queries and learnable category queries. Our method introduces image-guided query selection, which reduces the negative space by adaptively retrieving top-K relevant categories per image via cross-attention, thereby rebalancing gradient distributions and facilitating implicit hard example mining. Furthermore, CQ-DINO flexibly integrates explicit hierarchical category relationships in structured datasets (e.g., V3Det) or learns implicit category correlations via self-attention in generic datasets (e.g., COCO). Experiments demonstrate that CQ-DINO achieves superior performance on the challenging V3Det benchmark (surpassing previous methods by 2.1% AP) while maintaining competitiveness in COCO. Our work provides a scalable solution for real-world detection systems requiring wide category coverage. The code is publicly at https://github.com/FireRedTeam/CQ-DINO.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
Sun, Zhichao
Hu, Huazhang
Ma, Yidong
Liu, Gang
Chen, Yibo
Tang, Xu
Hu, Yao
Xu, Yongchao
Computer Vision and Pattern Recognition
With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive insufficient learning signals, and hard negative gradient dilution, where discriminative gradients are overwhelmed by numerous easy negatives. To address these challenges, we propose CQ-DINO, a category query-based object detection framework that reformulates classification as a contrastive task between object queries and learnable category queries. Our method introduces image-guided query selection, which reduces the negative space by adaptively retrieving top-K relevant categories per image via cross-attention, thereby rebalancing gradient distributions and facilitating implicit hard example mining. Furthermore, CQ-DINO flexibly integrates explicit hierarchical category relationships in structured datasets (e.g., V3Det) or learns implicit category correlations via self-attention in generic datasets (e.g., COCO). Experiments demonstrate that CQ-DINO achieves superior performance on the challenging V3Det benchmark (surpassing previous methods by 2.1% AP) while maintaining competitiveness in COCO. Our work provides a scalable solution for real-world detection systems requiring wide category coverage. The code is publicly at https://github.com/FireRedTeam/CQ-DINO.
title CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18430