DINO-Tok: Adapting DINO for Visual Tokenizers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Mingkai, Li, Mingxiao, Shu, Zhijian, Zheng, Anlin, Fan, Liaoyuan, Guo, Jiaxin, Shi, Tianxing, Lu, Dongyue, Li, Zeming, Guo, Xiaoyang, Qi, Xiaojuan, Long, Xiao-Xiao, Zhang, Qian, Tan, Ping, Yin, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
von: Fu, Weifu, et al.
Veröffentlicht: (2026)
von: Fu, Weifu, et al.
Veröffentlicht: (2026)
DINO-Foresight: Looking into the Future with DINO
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
von: Guo, Hao, et al.
Veröffentlicht: (2024)
von: Guo, Hao, et al.
Veröffentlicht: (2024)
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
von: Zimmermann, Robert, et al.
Veröffentlicht: (2026)
von: Zimmermann, Robert, et al.
Veröffentlicht: (2026)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
von: Ansari, Muhammad Musab
Veröffentlicht: (2025)
von: Ansari, Muhammad Musab
Veröffentlicht: (2025)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
von: Liu, Shilong, et al.
Veröffentlicht: (2023)
von: Liu, Shilong, et al.
Veröffentlicht: (2023)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
von: Gong, Ziren, et al.
Veröffentlicht: (2025)
von: Gong, Ziren, et al.
Veröffentlicht: (2025)
Text-guided Visual Prompt DINO for Generic Segmentation
von: Guan, Yuchen, et al.
Veröffentlicht: (2025)
von: Guan, Yuchen, et al.
Veröffentlicht: (2025)
DINO-AD: Unsupervised Anomaly Detection with Frozen DINO-V3 Features
von: Huo, Jiayu, et al.
Veröffentlicht: (2026)
von: Huo, Jiayu, et al.
Veröffentlicht: (2026)
Deploy DINO with Many-to-Many Association
von: Jiang, Haodong, et al.
Veröffentlicht: (2026)
von: Jiang, Haodong, et al.
Veröffentlicht: (2026)
2D Gaussians Meet Visual Tokenizer
von: Shi, Yiang, et al.
Veröffentlicht: (2025)
von: Shi, Yiang, et al.
Veröffentlicht: (2025)
DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video
von: Tumanyan, Narek, et al.
Veröffentlicht: (2024)
von: Tumanyan, Narek, et al.
Veröffentlicht: (2024)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
von: Zeng, Haoxi, et al.
Veröffentlicht: (2026)
von: Zeng, Haoxi, et al.
Veröffentlicht: (2026)
MammoDINO: Anatomically Aware Self-Supervision for Mammographic Images
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
PixelDINO: Semi-Supervised Semantic Segmentation for Detecting Permafrost Disturbances
von: Heidler, Konrad, et al.
Veröffentlicht: (2024)
von: Heidler, Konrad, et al.
Veröffentlicht: (2024)
Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection
von: Lu, Yehao, et al.
Veröffentlicht: (2025)
von: Lu, Yehao, et al.
Veröffentlicht: (2025)
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
von: Shu, Zhijian, et al.
Veröffentlicht: (2025)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
von: He, Xinwei, et al.
Veröffentlicht: (2026)
von: He, Xinwei, et al.
Veröffentlicht: (2026)
Unlocking Generalization in Polyp Segmentation with DINO Self-Attention "keys"
von: Monteiro, Carla, et al.
Veröffentlicht: (2025)
von: Monteiro, Carla, et al.
Veröffentlicht: (2025)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
von: Hu, Xiaotao, et al.
Veröffentlicht: (2024)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
von: Jiang, Dongsheng, et al.
Veröffentlicht: (2023)
von: Jiang, Dongsheng, et al.
Veröffentlicht: (2023)
Simplifying DINO via Coding Rate Regularization
von: Wu, Ziyang, et al.
Veröffentlicht: (2025)
von: Wu, Ziyang, et al.
Veröffentlicht: (2025)
DINO-VO: Learning Where to Focus for Enhanced State Estimation
von: Chen, Qi, et al.
Veröffentlicht: (2026)
von: Chen, Qi, et al.
Veröffentlicht: (2026)
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
von: Cui, Beilei, et al.
Veröffentlicht: (2024)
von: Cui, Beilei, et al.
Veröffentlicht: (2024)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
von: Azhari, Maulana Bisyir, et al.
Veröffentlicht: (2025)
von: Azhari, Maulana Bisyir, et al.
Veröffentlicht: (2025)
EndoDINO: A Foundation Model for GI Endoscopy
von: Dermyer, Patrick, et al.
Veröffentlicht: (2025)
von: Dermyer, Patrick, et al.
Veröffentlicht: (2025)
DIVE: Taming DINO for Subject-Driven Video Editing
von: Huang, Yi, et al.
Veröffentlicht: (2024)
von: Huang, Yi, et al.
Veröffentlicht: (2024)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
von: Singh, Rajhans, et al.
Veröffentlicht: (2025)
von: Singh, Rajhans, et al.
Veröffentlicht: (2025)
DINO as a von Mises-Fisher mixture model
von: Govindarajan, Hariprasath, et al.
Veröffentlicht: (2024)
von: Govindarajan, Hariprasath, et al.
Veröffentlicht: (2024)
FreqDINO: Frequency-Guided Adaptation for Generalized Boundary-Aware Ultrasound Image Segmentation
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
D$^3$FlowSLAM: Self-Supervised Dynamic SLAM with Flow Motion Decomposition and DINO Guidance
von: Yu, Xingyuan, et al.
Veröffentlicht: (2022)
von: Yu, Xingyuan, et al.
Veröffentlicht: (2022)
Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation
von: Han, Yizhao, et al.
Veröffentlicht: (2026)
von: Han, Yizhao, et al.
Veröffentlicht: (2026)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model
von: Nan, Zhixiong, et al.
Veröffentlicht: (2024)
von: Nan, Zhixiong, et al.
Veröffentlicht: (2024)
DINO-SD: Champion Solution for ICRA 2024 RoboDepth Challenge
von: Mao, Yifan, et al.
Veröffentlicht: (2024)
von: Mao, Yifan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization
von: Jia, Mingkai, et al.
Veröffentlicht: (2025) -
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
von: Fu, Weifu, et al.
Veröffentlicht: (2026) -
DINO-Foresight: Looking into the Future with DINO
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024) -
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
von: Guo, Hao, et al.
Veröffentlicht: (2024) -
DINO-QPM: Adapting Visual Foundation Models for Globally Interpretable Image Classification
von: Zimmermann, Robert, et al.
Veröffentlicht: (2026)