Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Csizmadia, Daniel, Codreanu, Andrei, Sim, Victor, Prabhu, Vighnesh, Lu, Michael, Zhu, Kevin, O'Brien, Sean, Sharma, Vasu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026)
by: Chung, Jeannie, et al.
Published: (2026)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
by: Do, Timothy, et al.
Published: (2025)
by: Do, Timothy, et al.
Published: (2025)
Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
by: Said, Muna Numan, et al.
Published: (2025)
by: Said, Muna Numan, et al.
Published: (2025)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
by: Chou, Cheng-Ting, et al.
Published: (2025)
by: Chou, Cheng-Ting, et al.
Published: (2025)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
by: Jo, Sanghyun, et al.
Published: (2024)
by: Jo, Sanghyun, et al.
Published: (2024)
Choosing Wisely and Learning Deeply: Selective Cross-Modality Distillation via CLIP for Domain Generalization
by: Leng, Jixuan, et al.
Published: (2023)
by: Leng, Jixuan, et al.
Published: (2023)
Rewrite-to-Rank: Optimizing Ad Visibility via Retrieval-Aware Text Rewriting
by: Ho, Chloe, et al.
Published: (2025)
by: Ho, Chloe, et al.
Published: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
by: Xiong, Lang, et al.
Published: (2025)
by: Xiong, Lang, et al.
Published: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
by: Begin, James, et al.
Published: (2025)
by: Begin, James, et al.
Published: (2025)
ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
by: Khalid, Haziq Mohammad, et al.
Published: (2025)
by: Khalid, Haziq Mohammad, et al.
Published: (2025)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
by: Le, Duy, et al.
Published: (2025)
by: Le, Duy, et al.
Published: (2025)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
by: Kao, Kuei-Chun
Published: (2024)
by: Kao, Kuei-Chun
Published: (2024)
Graph-Based Cross-Domain Knowledge Distillation for Cross-Dataset Text-to-Image Person Retrieval
by: Luo, Bingjun, et al.
Published: (2025)
by: Luo, Bingjun, et al.
Published: (2025)
SMAGDi: Socratic Multi Agent Interaction Graph Distillation for Efficient High Accuracy Reasoning
by: Aluru, Aayush, et al.
Published: (2025)
by: Aluru, Aayush, et al.
Published: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
by: Liu, Joshua, et al.
Published: (2025)
by: Liu, Joshua, et al.
Published: (2025)
Translate-Distill: Learning Cross-Language Dense Retrieval by Translation and Distillation
by: Yang, Eugene, et al.
Published: (2024)
by: Yang, Eugene, et al.
Published: (2024)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
by: Baek, Shaun, et al.
Published: (2025)
by: Baek, Shaun, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
by: Boizard, Nicolas, et al.
Published: (2024)
by: Boizard, Nicolas, et al.
Published: (2024)
CLIP-KD: An Empirical Study of CLIP Model Distillation
by: Yang, Chuanguang, et al.
Published: (2023)
by: Yang, Chuanguang, et al.
Published: (2023)
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
by: Yu, Stanley, et al.
Published: (2025)
by: Yu, Stanley, et al.
Published: (2025)
MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact
by: Nandan, Advey, et al.
Published: (2025)
by: Nandan, Advey, et al.
Published: (2025)
Distillation Enhanced Generative Retrieval
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Channel Attention-Guided Cross-Modal Knowledge Distillation for Referring Image Segmentation
by: Yang, Chen
Published: (2026)
by: Yang, Chen
Published: (2026)
SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
by: Buess, Lukas, et al.
Published: (2025)
by: Buess, Lukas, et al.
Published: (2025)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
by: Yugeswardeenoo, Dharunish, et al.
Published: (2024)
Enriching Knowledge Distillation with Cross-Modal Teacher Fusion
by: Mansourian, Amir M., et al.
Published: (2025)
by: Mansourian, Amir M., et al.
Published: (2025)
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
by: Yin, Zhichao, et al.
Published: (2024)
by: Yin, Zhichao, et al.
Published: (2024)
Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
by: Weng, Yuzhe, et al.
Published: (2024)
by: Weng, Yuzhe, et al.
Published: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
by: Che, Chang, et al.
Published: (2024)
by: Che, Chang, et al.
Published: (2024)
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation
by: Zhu, Xiaofei, et al.
Published: (2024)
by: Zhu, Xiaofei, et al.
Published: (2024)
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
ClinKD: Cross-Modal Clinical Knowledge Distiller For Multi-Task Medical Images
by: Ge, Hongyu, et al.
Published: (2025)
by: Ge, Hongyu, et al.
Published: (2025)
CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation
by: Wang, Wenxuan, et al.
Published: (2023)
by: Wang, Wenxuan, et al.
Published: (2023)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
Similar Items
-
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026) -
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
by: Gupta, Abhay, et al.
Published: (2025) -
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
by: Do, Timothy, et al.
Published: (2025) -
Deconstructing Bias: A Multifaceted Framework for Diagnosing Cultural and Compositional Inequities in Text-to-Image Generative Models
by: Said, Muna Numan, et al.
Published: (2025) -
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)