Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shilong, Zeng, Zhaoyang, Ren, Tianhe, Li, Feng, Zhang, Hao, Yang, Jie, Jiang, Qing, Li, Chunyuan, Yang, Jianwei, Su, Hang, Zhu, Jun, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
by: Qu, Jinyuan, et al.
Published: (2025)
by: Qu, Jinyuan, et al.
Published: (2025)
T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
by: Jiang, Qing, et al.
Published: (2024)
by: Jiang, Qing, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
by: Li, Aiden Yiliu, et al.
Published: (2025)
by: Li, Aiden Yiliu, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
by: Singh, Rajhans, et al.
Published: (2025)
by: Singh, Rajhans, et al.
Published: (2025)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
by: Gong, Ziren, et al.
Published: (2025)
by: Gong, Ziren, et al.
Published: (2025)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
Simplifying DINO via Coding Rate Regularization
by: Wu, Ziyang, et al.
Published: (2025)
by: Wu, Ziyang, et al.
Published: (2025)
TAPTR: Tracking Any Point with Transformers as Detection
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long Video
by: Qu, Jinyuan, et al.
Published: (2024)
by: Qu, Jinyuan, et al.
Published: (2024)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
by: Dulal, Rabin, et al.
Published: (2025)
by: Dulal, Rabin, et al.
Published: (2025)
TAPTRv2: Attention-based Position Update Improves Tracking Any Point
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization
by: Wang, Zanyi, et al.
Published: (2026)
by: Wang, Zanyi, et al.
Published: (2026)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attention Network for T1-to-BOLD Generation
by: Wang, Jianwei, et al.
Published: (2025)
by: Wang, Jianwei, et al.
Published: (2025)
Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App
by: Nossair, Abdelilah, et al.
Published: (2024)
by: Nossair, Abdelilah, et al.
Published: (2024)
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
by: He, Xinwei, et al.
Published: (2026)
by: He, Xinwei, et al.
Published: (2026)
DINO-AD: Unsupervised Anomaly Detection with Frozen DINO-V3 Features
by: Huo, Jiayu, et al.
Published: (2026)
by: Huo, Jiayu, et al.
Published: (2026)
Deploy DINO with Many-to-Many Association
by: Jiang, Haodong, et al.
Published: (2026)
by: Jiang, Haodong, et al.
Published: (2026)
Efficient License Plate Recognition via Pseudo-Labeled Supervision with Grounding DINO and YOLOv8
by: Vargoorani, Zahra Ebrahimi, et al.
Published: (2025)
by: Vargoorani, Zahra Ebrahimi, et al.
Published: (2025)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video
by: Tumanyan, Narek, et al.
Published: (2024)
by: Tumanyan, Narek, et al.
Published: (2024)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
by: Nan, Zhixiong, et al.
Published: (2024)
by: Nan, Zhixiong, et al.
Published: (2024)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
by: Liang, Zhuonan, et al.
Published: (2026)
by: Liang, Zhuonan, et al.
Published: (2026)
Multi-task Image Restoration Guided By Robust DINO Features
by: Lin, Xin, et al.
Published: (2023)
by: Lin, Xin, et al.
Published: (2023)
Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
by: Chen, Zhiyang, et al.
Published: (2025)
by: Chen, Zhiyang, et al.
Published: (2025)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)
by: Juneja, Shubham, et al.
Published: (2024)
DINO-YOLO: Self-Supervised Pre-training for Data-Efficient Object Detection in Civil Engineering Applications
by: P, Malaisree, et al.
Published: (2025)
by: P, Malaisree, et al.
Published: (2025)
Similar Items
-
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024) -
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024) -
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026) -
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025) -
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)