DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Shuyang, Yu, Nan, Zhang, Yiming, Ding, Zenghui, Wu, Zhenyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024)
by: Luo, Lin, et al.
Published: (2024)
Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
by: Yang, Haiyu, et al.
Published: (2026)
by: Yang, Haiyu, et al.
Published: (2026)
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025)
by: Pham, Phuc, et al.
Published: (2025)
M$^3$HL: Mutual Mask Mix with High-Low Level Feature Consistency for Semi-Supervised Medical Image Segmentation
by: Liu, Yajun, et al.
Published: (2025)
by: Liu, Yajun, et al.
Published: (2025)
Changes in Gaza: DINOv3-Powered Multi-Class Change Detection for Damage Assessment in Conflict Zones
by: Zheng, Kai, et al.
Published: (2025)
by: Zheng, Kai, et al.
Published: (2025)
Adversarial Prompt Injection Attack on Multimodal Large Language Models
by: Ding, Meiwen, et al.
Published: (2026)
by: Ding, Meiwen, et al.
Published: (2026)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
by: Liu, Yufang, et al.
Published: (2024)
by: Liu, Yufang, et al.
Published: (2024)
DINOv3 with Test-Time Training for Medical Image Registration
by: Wang, Shansong, et al.
Published: (2025)
by: Wang, Shansong, et al.
Published: (2025)
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach
by: Guan, Jiwei, et al.
Published: (2024)
by: Guan, Jiwei, et al.
Published: (2024)
TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
by: Zhang, Shu-Hao, et al.
Published: (2025)
by: Zhang, Shu-Hao, et al.
Published: (2025)
DINOv3 Meets YOLO26 for Weed Detection in Vegetable Crops
by: Deng, Boyang, et al.
Published: (2026)
by: Deng, Boyang, et al.
Published: (2026)
SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
by: Lu, Wenbo
Published: (2025)
by: Lu, Wenbo
Published: (2025)
Evolving Prompt Adaptation for Vision-Language Models
by: Zhang, Enming, et al.
Published: (2026)
by: Zhang, Enming, et al.
Published: (2026)
DINO-MVR: Multi-View Readout of Frozen DINOv3 for Annotation-Efficient Medical Segmentation
by: Jiang, Wei, et al.
Published: (2026)
by: Jiang, Wei, et al.
Published: (2026)
DINOv3 as a Frozen Encoder for CRPS-Oriented Probabilistic Rainfall Nowcasting
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
Physical Prompt Injection Attacks on Large Vision-Language Models
by: Ling, Chen, et al.
Published: (2026)
by: Ling, Chen, et al.
Published: (2026)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
by: Mei, Guofeng, et al.
Published: (2024)
by: Mei, Guofeng, et al.
Published: (2024)
Evaluating General Purpose Vision Foundation Models for Medical Image Analysis: An Experimental Study of DINOv2 on Radiology Benchmarks
by: Baharoon, Mohammed, et al.
Published: (2023)
by: Baharoon, Mohammed, et al.
Published: (2023)
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
by: Saini, Harshvardhan, et al.
Published: (2026)
by: Saini, Harshvardhan, et al.
Published: (2026)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
Characterizing Disparity Between Edge Models and High-Accuracy Base Models for Vision Tasks
by: Wang, Zhenyu, et al.
Published: (2024)
by: Wang, Zhenyu, et al.
Published: (2024)
Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
by: Chen, Minglei, et al.
Published: (2026)
by: Chen, Minglei, et al.
Published: (2026)
Revealing the Semantic Selection Gap in DINOv3 through Training-Free Few-Shot Segmentation
by: Zakir, Hussni Mohd, et al.
Published: (2026)
by: Zakir, Hussni Mohd, et al.
Published: (2026)
Multimodal Distribution Matching for Vision-Language Dataset Distillation
by: Jeong, Jongoh, et al.
Published: (2026)
by: Jeong, Jongoh, et al.
Published: (2026)
Reward Guided Latent Consistency Distillation
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
by: Song, Chull Hwan, et al.
Published: (2024)
by: Song, Chull Hwan, et al.
Published: (2024)
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
by: Kumar, Ashwin, et al.
Published: (2026)
by: Kumar, Ashwin, et al.
Published: (2026)
Denoising Score Distillation: From Noisy Diffusion Pretraining to One-Step High-Quality Generation
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
Distilling 3D Spatial Reasoning into a Lightweight Vision-Language Model with CoT
by: Asfour, Alaa, et al.
Published: (2026)
by: Asfour, Alaa, et al.
Published: (2026)
Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
by: Ortega, Joel Valdivia, et al.
Published: (2025)
by: Ortega, Joel Valdivia, et al.
Published: (2025)
General Purpose Image Encoder DINOv2 for Medical Image Registration
by: Song, Xinrui, et al.
Published: (2024)
by: Song, Xinrui, et al.
Published: (2024)
Homogeneous and Heterogeneous Consistency progressive Re-ranking for Visible-Infrared Person Re-identification
by: Wang, Yiming
Published: (2026)
by: Wang, Yiming
Published: (2026)
DINOv3
by: Siméoni, Oriane, et al.
Published: (2025)
by: Siméoni, Oriane, et al.
Published: (2025)
Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
Similar Items
-
RankCLIP: Ranking-Consistent Language-Image Pretraining
by: Zhang, Yiming, et al.
Published: (2024) -
Adversarial Prompt Distillation for Vision-Language Models
by: Luo, Lin, et al.
Published: (2024) -
Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
by: Yang, Haiyu, et al.
Published: (2026) -
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
by: Pham, Phuc, et al.
Published: (2025) -
M$^3$HL: Mutual Mask Mix with High-Low Level Feature Consistency for Semi-Supervised Medical Image Segmentation
by: Liu, Yajun, et al.
Published: (2025)