DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xinwei, Zheng, Yansong, Han, Qianru, Wang, Zhichuan, Cai, Yuxuan, Zhou, Yang, Xia, Jingbo, Wang, Yulong, Xiang, Jinhai, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
by: Wang, Zhichuan, et al.
Published: (2025)
by: Wang, Zhichuan, et al.
Published: (2025)
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
by: Han, Qianru, et al.
Published: (2024)
by: Han, Qianru, et al.
Published: (2024)
Tetrahedron-Net for Medical Image Registration
by: Xiang, Jinhai, et al.
Published: (2025)
by: Xiang, Jinhai, et al.
Published: (2025)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
Not All Regions Are Equal: Attention-Guided Perturbation Network for Industrial Anomaly Detection
by: Huang, Tingfeng, et al.
Published: (2024)
by: Huang, Tingfeng, et al.
Published: (2024)
SimpleFusion: A Simple Fusion Framework for Infrared and Visible Images
by: Chen, Ming, et al.
Published: (2024)
by: Chen, Ming, et al.
Published: (2024)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
by: Bai, Sule, et al.
Published: (2024)
by: Bai, Sule, et al.
Published: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
by: Nan, Zhixiong, et al.
Published: (2024)
by: Nan, Zhixiong, et al.
Published: (2024)
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
by: Zhu, Wenqi, et al.
Published: (2024)
by: Zhu, Wenqi, et al.
Published: (2024)
Online Open-set Semi-supervised Object Detection with Dual Competing Head
by: Wang, Zerun, et al.
Published: (2023)
by: Wang, Zerun, et al.
Published: (2023)
To Eat or Not to Eat
by: Altmann, Peter, et al.
Published: (2024)
by: Altmann, Peter, et al.
Published: (2024)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Omni-AD: Learning to Reconstruct Global and Local Features for Multi-class Anomaly Detection
by: Quan, Jiajie, et al.
Published: (2025)
by: Quan, Jiajie, et al.
Published: (2025)
Beyond Known Objects: A Novel Framework for Open-Set Object Detection using Negative-Aware Norm
by: Zhang, Yuchen, et al.
Published: (2026)
by: Zhang, Yuchen, et al.
Published: (2026)
Known Meets Unknown: Mitigating Overconfidence in Open Set Recognition
by: Zhao, Dongdong, et al.
Published: (2025)
by: Zhao, Dongdong, et al.
Published: (2025)
Looking Beyond the Known: Towards a Data Discovery Guided Open-World Object Detection
by: Majee, Anay, et al.
Published: (2025)
by: Majee, Anay, et al.
Published: (2025)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)
by: Wysoczańska, Monika, et al.
Published: (2023)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
by: Lu, Yehao, et al.
Published: (2025)
by: Lu, Yehao, et al.
Published: (2025)
Open-Set Object Detection By Aligning Known Class Representations
by: Sarkar, Hiran, et al.
Published: (2024)
by: Sarkar, Hiran, et al.
Published: (2024)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
by: Asgarov, Ali, et al.
Published: (2024)
by: Asgarov, Ali, et al.
Published: (2024)
Beyond the Known: Investigating LLMs Performance on Out-of-Domain Intent Detection
by: Wang, Pei, et al.
Published: (2024)
by: Wang, Pei, et al.
Published: (2024)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
by: Xiao, Linhui, et al.
Published: (2023)
by: Xiao, Linhui, et al.
Published: (2023)
LLM4SG: Adapting Large Language Model for Scatterer Generation via Synesthesia of Machines
by: Han, Zengrui, et al.
Published: (2025)
by: Han, Zengrui, et al.
Published: (2025)
DODA: Adapting Object Detectors to Dynamic Agricultural Environments in Real-Time with Diffusion
by: Xiang, Shuai, et al.
Published: (2024)
by: Xiang, Shuai, et al.
Published: (2024)
Efficiently Disentangling CLIP for Multi-Object Perception
by: Rawlekar, Samyak, et al.
Published: (2025)
by: Rawlekar, Samyak, et al.
Published: (2025)
Beyond the Known: Enhancing Open Set Domain Adaptation with Unknown Exploration
by: Silva, Lucas Fernando Alvarenga e, et al.
Published: (2024)
by: Silva, Lucas Fernando Alvarenga e, et al.
Published: (2024)
Beyond the Known: Novel Class Discovery for Open-world Graph Learning
by: Jin, Yucheng, et al.
Published: (2024)
by: Jin, Yucheng, et al.
Published: (2024)
LLM4MG: Adapting Large Language Model for Multipath Generation via Synesthesia of Machines
by: Huang, Ziwei, et al.
Published: (2025)
by: Huang, Ziwei, et al.
Published: (2025)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
by: Cao, Guiping, et al.
Published: (2025)
by: Cao, Guiping, et al.
Published: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
by: Zimmermann, Robert, et al.
Published: (2026)
by: Zimmermann, Robert, et al.
Published: (2026)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
by: Jiang, Dongsheng, et al.
Published: (2023)
by: Jiang, Dongsheng, et al.
Published: (2023)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
by: Imam, Mohamed Fazli, et al.
Published: (2024)
by: Imam, Mohamed Fazli, et al.
Published: (2024)
Similar Items
-
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
by: Wang, Zhichuan, et al.
Published: (2025) -
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
by: Wang, Zhichuan, et al.
Published: (2025) -
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
by: Han, Qianru, et al.
Published: (2024) -
Tetrahedron-Net for Medical Image Registration
by: Xiang, Jinhai, et al.
Published: (2025) -
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)