When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Dulal, Rabin, Zheng, Lihong, Kabir, Muhammad Ashad |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CCoMAML: Efficient Cattle Identification Using Cooperative Model-Agnostic Meta-Learning
by: Dulal, Rabin, et al.
Published: (2025)
by: Dulal, Rabin, et al.
Published: (2025)
MHAFF: Multi-Head Attention Feature Fusion of CNN and Transformer for Cattle Identification
by: Dulal, Rabin, et al.
Published: (2025)
by: Dulal, Rabin, et al.
Published: (2025)
WoundFormer: Multi-Scale Spatial Feature Fusion for Multi-Class Wound Tissue Segmentation
by: Kabir, Muhammad Ashad, et al.
Published: (2026)
by: Kabir, Muhammad Ashad, et al.
Published: (2026)
Agreement-Driven Multi-View 3D Reconstruction for Live Cattle Weight Estimation
by: Dulal, Rabin, et al.
Published: (2026)
by: Dulal, Rabin, et al.
Published: (2026)
Brain Tumor Identification using Improved YOLOv8
by: Dulal, Rupesh, et al.
Published: (2025)
by: Dulal, Rupesh, et al.
Published: (2025)
Muzzle-Based Cattle Identification System Using Artificial Intelligence (AI)
by: Islam, Hasan Zohirul, et al.
Published: (2024)
by: Islam, Hasan Zohirul, et al.
Published: (2024)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types
by: Hanum, Sauda Adiv, et al.
Published: (2025)
by: Hanum, Sauda Adiv, et al.
Published: (2025)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
From Lab to Pocket: A Novel Continual Learning-based Mobile Application for Screening COVID-19
by: Falero, Danny, et al.
Published: (2024)
by: Falero, Danny, et al.
Published: (2024)
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
by: Chen, Lihong, et al.
Published: (2025)
by: Chen, Lihong, et al.
Published: (2025)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
by: He, Xunjie, et al.
Published: (2025)
by: He, Xunjie, et al.
Published: (2025)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
by: Singh, Rajhans, et al.
Published: (2025)
by: Singh, Rajhans, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
by: Islam, Md Touhidul, et al.
Published: (2025)
by: Islam, Md Touhidul, et al.
Published: (2025)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
Grounding DINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models
by: Rasaee, Hamza, et al.
Published: (2025)
by: Rasaee, Hamza, et al.
Published: (2025)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
by: Liang, Zhuonan, et al.
Published: (2026)
by: Liang, Zhuonan, et al.
Published: (2026)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
by: Glória-Silva, Diogo, et al.
Published: (2024)
by: Glória-Silva, Diogo, et al.
Published: (2024)
TuneVLSeg: Prompt Tuning Benchmark for Vision-Language Segmentation Models
by: Adhikari, Rabin, et al.
Published: (2024)
by: Adhikari, Rabin, et al.
Published: (2024)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
by: Kim, Jihyeok, et al.
Published: (2025)
by: Kim, Jihyeok, et al.
Published: (2025)
Multi-task Image Restoration Guided By Robust DINO Features
by: Lin, Xin, et al.
Published: (2023)
by: Lin, Xin, et al.
Published: (2023)
ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
by: Hassan, Md Zahim, et al.
Published: (2025)
by: Hassan, Md Zahim, et al.
Published: (2025)
SpectraDINO: Bridging the Spectral Gap in Vision Foundation Models via Lightweight Adapters
by: Nalcakan, Yagiz, et al.
Published: (2026)
by: Nalcakan, Yagiz, et al.
Published: (2026)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
by: Nan, Zhixiong, et al.
Published: (2024)
by: Nan, Zhixiong, et al.
Published: (2024)
Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App
by: Nossair, Abdelilah, et al.
Published: (2024)
by: Nossair, Abdelilah, et al.
Published: (2024)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
by: Pan, Chenbin, et al.
Published: (2025)
by: Pan, Chenbin, et al.
Published: (2025)
When and Where to Attack? Stage-wise Attention-Guided Adversarial Attack on Large Vision Language Models
by: Kwak, Jaehyun, et al.
Published: (2026)
by: Kwak, Jaehyun, et al.
Published: (2026)
ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
by: Hu, Junyi, et al.
Published: (2026)
by: Hu, Junyi, et al.
Published: (2026)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
by: Gong, Ziren, et al.
Published: (2025)
by: Gong, Ziren, et al.
Published: (2025)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video
by: Tumanyan, Narek, et al.
Published: (2024)
by: Tumanyan, Narek, et al.
Published: (2024)
Similar Items
-
CCoMAML: Efficient Cattle Identification Using Cooperative Model-Agnostic Meta-Learning
by: Dulal, Rabin, et al.
Published: (2025) -
MHAFF: Multi-Head Attention Feature Fusion of CNN and Transformer for Cattle Identification
by: Dulal, Rabin, et al.
Published: (2025) -
WoundFormer: Multi-Scale Spatial Feature Fusion for Multi-Class Wound Tissue Segmentation
by: Kabir, Muhammad Ashad, et al.
Published: (2026) -
Agreement-Driven Multi-View 3D Reconstruction for Live Cattle Weight Estimation
by: Dulal, Rabin, et al.
Published: (2026) -
Brain Tumor Identification using Improved YOLOv8
by: Dulal, Rupesh, et al.
Published: (2025)