Saved in:
| Main Authors: | Wu, Ziyang, Zhang, Jingyuan, Pai, Druv, Wang, XuDong, Singh, Chandan, Yang, Jianwei, Gao, Jianfeng, Ma, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.10385 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023)
by: Yu, Yaodong, et al.
Published: (2023)
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
DINO-AD: Unsupervised Anomaly Detection with Frozen DINO-V3 Features
by: Huo, Jiayu, et al.
Published: (2026)
by: Huo, Jiayu, et al.
Published: (2026)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
Towards Understanding Graphical Perception in Large Multimodal Models
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Pix2Gif: Motion-Guided Diffusion for GIF Generation
by: Kandala, Hitesh, et al.
Published: (2024)
by: Kandala, Hitesh, et al.
Published: (2024)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attention Network for T1-to-BOLD Generation
by: Wang, Jianwei, et al.
Published: (2025)
by: Wang, Jianwei, et al.
Published: (2025)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
OVS-DINO: Open-Vocabulary Segmentation via Structure-Aligned SAM-DINO with Language Guidance
by: Zeng, Haoxi, et al.
Published: (2026)
by: Zeng, Haoxi, et al.
Published: (2026)
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
by: Jain, Jitesh, et al.
Published: (2024)
by: Jain, Jitesh, et al.
Published: (2024)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
by: Gong, Ziren, et al.
Published: (2025)
by: Gong, Ziren, et al.
Published: (2025)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection
by: Sun, Zhichao, et al.
Published: (2025)
by: Sun, Zhichao, et al.
Published: (2025)
Deploy DINO with Many-to-Many Association
by: Jiang, Haodong, et al.
Published: (2026)
by: Jiang, Haodong, et al.
Published: (2026)
Few-Shot Adaptation of Grounding DINO for Agricultural Domain
by: Singh, Rajhans, et al.
Published: (2025)
by: Singh, Rajhans, et al.
Published: (2025)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Deformable One-shot Face Stylization via DINO Semantic Guidance
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
DINO-VO: Learning Where to Focus for Enhanced State Estimation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
DINO-Explorer: Active Underwater Discovery via Ego-Motion Compensated Semantic Predictive Coding
by: Jin, Yuhan, et al.
Published: (2026)
by: Jin, Yuhan, et al.
Published: (2026)
Language-Image Alignment with Fixed Text Encoders
by: Yang, Jingfeng, et al.
Published: (2025)
by: Yang, Jingfeng, et al.
Published: (2025)
Two-stage Rainfall-Forecasting Diffusion Model
by: Ling, XuDong, et al.
Published: (2024)
by: Ling, XuDong, et al.
Published: (2024)
DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video
by: Tumanyan, Narek, et al.
Published: (2024)
by: Tumanyan, Narek, et al.
Published: (2024)
Human detectors are surprisingly powerful reward models
by: Ashutosh, Kumar, et al.
Published: (2026)
by: Ashutosh, Kumar, et al.
Published: (2026)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
by: Meng, Lingchen, et al.
Published: (2024)
by: Meng, Lingchen, et al.
Published: (2024)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Representation Alignment Contrastive Regularization for Multi-Object Tracking
by: Liu, Zhonglin, et al.
Published: (2024)
by: Liu, Zhonglin, et al.
Published: (2024)
Fine-Grained DINO Tuning with Dual Supervision for Face Forgery Detection
by: Zhang, Tianxiang, et al.
Published: (2025)
by: Zhang, Tianxiang, et al.
Published: (2025)
Tri-path DINO: Feature Complementary Learning for Remote Sensing Multi-Class Change Detection
by: Zheng, Kai, et al.
Published: (2026)
by: Zheng, Kai, et al.
Published: (2026)
Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter
by: Chen, Zhiyang, et al.
Published: (2025)
by: Chen, Zhiyang, et al.
Published: (2025)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
by: Ren, Tianhe, et al.
Published: (2024)
by: Ren, Tianhe, et al.
Published: (2024)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
DIVE: Taming DINO for Subject-Driven Video Editing
by: Huang, Yi, et al.
Published: (2024)
by: Huang, Yi, et al.
Published: (2024)
Similar Items
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is?
by: Yu, Yaodong, et al.
Published: (2023) -
Scaling White-Box Transformers for Vision
by: Yang, Jinrui, et al.
Published: (2024) -
DINO-AD: Unsupervised Anomaly Detection with Frozen DINO-V3 Features
by: Huo, Jiayu, et al.
Published: (2026) -
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024) -
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)