CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Zhichao, Hu, Huazhang, Ma, Yidong, Liu, Gang, Chen, Yibo, Tang, Xu, Hu, Yao, Xu, Yongchao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
by: Sun, Zhichao, et al.
Published: (2026)
by: Sun, Zhichao, et al.
Published: (2026)
VastTrack: Vast Category Visual Object Tracking
by: Peng, Liang, et al.
Published: (2024)
by: Peng, Liang, et al.
Published: (2024)
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
by: Chen, Yitong, et al.
Published: (2024)
by: Chen, Yitong, et al.
Published: (2024)
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
by: Wang, Jiaqi, et al.
Published: (2024)
by: Wang, Jiaqi, et al.
Published: (2024)
Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024
by: Wu, Peixi, et al.
Published: (2024)
by: Wu, Peixi, et al.
Published: (2024)
VastGaussian: Vast 3D Gaussians for Large Scene Reconstruction
by: Lin, Jiaqi, et al.
Published: (2024)
by: Lin, Jiaqi, et al.
Published: (2024)
Category Query Learning for Human-Object Interaction Classification
by: Xie, Chi, et al.
Published: (2023)
by: Xie, Chi, et al.
Published: (2023)
Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
by: Hu, Yupeng, et al.
Published: (2025)
by: Hu, Yupeng, et al.
Published: (2025)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
Shifted Autoencoders for Point Annotation Restoration in Object Counting
by: Zou, Yuda, et al.
Published: (2023)
by: Zou, Yuda, et al.
Published: (2023)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
by: Zhang, Yupeng, et al.
Published: (2026)
by: Zhang, Yupeng, et al.
Published: (2026)
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
by: Yao, Jianhang, et al.
Published: (2025)
by: Yao, Jianhang, et al.
Published: (2025)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Mitigating the Impact of Prominent Position Shift in Drone-based RGBT Object Detection
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Position-Guided Prompt Learning for Anomaly Detection in Chest X-Rays
by: Sun, Zhichao, et al.
Published: (2024)
by: Sun, Zhichao, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection
by: Zhao, Kaixiang, et al.
Published: (2026)
by: Zhao, Kaixiang, et al.
Published: (2026)
Simplifying DINO via Coding Rate Regularization
by: Wu, Ziyang, et al.
Published: (2025)
by: Wu, Ziyang, et al.
Published: (2025)
Boosting Open-Vocabulary Object Detection by Handling Background Samples
by: Zeng, Ruizhe, et al.
Published: (2024)
by: Zeng, Ruizhe, et al.
Published: (2024)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
Mitigating Query Selection Bias in Referring Video Object Segmentation
by: Zhang, Dingwei, et al.
Published: (2025)
by: Zhang, Dingwei, et al.
Published: (2025)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
Open Vocabulary Monocular 3D Object Detection
by: Yao, Jin, et al.
Published: (2024)
by: Yao, Jin, et al.
Published: (2024)
Collaborative Novel Object Discovery and Box-Guided Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects
by: Wang, Jiawei, et al.
Published: (2025)
by: Wang, Jiawei, et al.
Published: (2025)
HA-FGOVD: Highlighting Fine-grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
by: Ma, Yuqi, et al.
Published: (2024)
by: Ma, Yuqi, et al.
Published: (2024)
Streamlined Open-Vocabulary Human-Object Interaction Detection
by: Sun, Chang, et al.
Published: (2026)
by: Sun, Chang, et al.
Published: (2026)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
by: Wang, Siheng, et al.
Published: (2026)
by: Wang, Siheng, et al.
Published: (2026)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
by: Nan, Zhixiong, et al.
Published: (2024)
by: Nan, Zhixiong, et al.
Published: (2024)
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
by: Li, Jiaming, et al.
Published: (2026)
by: Li, Jiaming, et al.
Published: (2026)
DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery
by: Faulkenberry, Ryan, et al.
Published: (2026)
by: Faulkenberry, Ryan, et al.
Published: (2026)
Locating and Mitigating Gradient Conflicts in Point Cloud Domain Adaptation via Saliency Map Skewness
by: Tang, Jiaqi, et al.
Published: (2025)
by: Tang, Jiaqi, et al.
Published: (2025)
OVGrasp: Open-Vocabulary Grasping Assistance via Multimodal Intent Detection
by: Hu, Chen, et al.
Published: (2025)
by: Hu, Chen, et al.
Published: (2025)
Scaling Open-Vocabulary Object Detection
by: Minderer, Matthias, et al.
Published: (2023)
by: Minderer, Matthias, et al.
Published: (2023)
Similar Items
-
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
by: Sun, Zhichao, et al.
Published: (2026) -
VastTrack: Vast Category Visual Object Tracking
by: Peng, Liang, et al.
Published: (2024) -
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection
by: Chen, Yitong, et al.
Published: (2024) -
V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
by: Wang, Jiaqi, et al.
Published: (2024) -
Enhanced Object Detection: A Study on Vast Vocabulary Object Detection Track for V3Det Challenge 2024
by: Wu, Peixi, et al.
Published: (2024)