DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Zhichao, Hui, Binyuan, Yang, Min, Huang, Fei, Li, Yongbin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
by: Lang, Hao, et al.
Published: (2023)
by: Lang, Hao, et al.
Published: (2023)
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
by: Cai, Yucheng, et al.
Published: (2024)
by: Cai, Yucheng, et al.
Published: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
Iterative Forward Tuning Boosts In-Context Learning in Language Models
by: Yang, Jiaxi, et al.
Published: (2023)
by: Yang, Jiaxi, et al.
Published: (2023)
Multi-dimensional Evaluation of Empathetic Dialog Responses
by: Xu, Zhichao, et al.
Published: (2024)
by: Xu, Zhichao, et al.
Published: (2024)
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
by: Rabbani, Parisa, et al.
Published: (2026)
by: Rabbani, Parisa, et al.
Published: (2026)
A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
by: Zhao, Yingxiu, et al.
Published: (2023)
by: Zhao, Yingxiu, et al.
Published: (2023)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
by: Chew, Oscar, et al.
Published: (2026)
by: Chew, Oscar, et al.
Published: (2026)
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)
by: Faghri, Fartash, et al.
Published: (2025)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
by: Asgarov, Ali, et al.
Published: (2024)
by: Asgarov, Ali, et al.
Published: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
by: Xie, Jingyou, et al.
Published: (2024)
by: Xie, Jingyou, et al.
Published: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
by: Chen, Junjie, et al.
Published: (2024)
by: Chen, Junjie, et al.
Published: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
Demystifying CLIP Data
by: Xu, Hu, et al.
Published: (2023)
by: Xu, Hu, et al.
Published: (2023)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
by: Koukounas, Andreas, et al.
Published: (2024)
by: Koukounas, Andreas, et al.
Published: (2024)
LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
by: Huang, Weiquan, et al.
Published: (2024)
by: Huang, Weiquan, et al.
Published: (2024)
TiC-CLIP: Continual Training of CLIP Models
by: Garg, Saurabh, et al.
Published: (2023)
by: Garg, Saurabh, et al.
Published: (2023)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
by: Vasu, Pavan Kumar Anasosalu, et al.
Published: (2023)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
by: Wang, Enguang, et al.
Published: (2024)
by: Wang, Enguang, et al.
Published: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
by: Shah, Siddhant Bikram, et al.
Published: (2024)
by: Shah, Siddhant Bikram, et al.
Published: (2024)
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
by: Zhang, Xinghua, et al.
Published: (2024)
by: Zhang, Xinghua, et al.
Published: (2024)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
Meta CLIP 2: A Worldwide Scaling Recipe
by: Chuang, Yung-Sung, et al.
Published: (2025)
by: Chuang, Yung-Sung, et al.
Published: (2025)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
by: Eslami, Sedigheh, et al.
Published: (2024)
by: Eslami, Sedigheh, et al.
Published: (2024)
One-Shot Learning as Instruction Data Prospector for Large Language Models
by: Li, Yunshui, et al.
Published: (2023)
by: Li, Yunshui, et al.
Published: (2023)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
by: Lin, Haokun, et al.
Published: (2025)
by: Lin, Haokun, et al.
Published: (2025)
Entriever: Energy-based Retriever for Knowledge-Grounded Dialog Systems
by: Cai, Yucheng, et al.
Published: (2025)
by: Cai, Yucheng, et al.
Published: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
by: Wang, Hsuan-Fu, et al.
Published: (2024)
by: Wang, Hsuan-Fu, et al.
Published: (2024)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
by: Wu, Ruijia, et al.
Published: (2025)
by: Wu, Ruijia, et al.
Published: (2025)
Language-Conditioned Visual Grounding with CLIP Multilingual
by: de Curtò, J., et al.
Published: (2026)
by: de Curtò, J., et al.
Published: (2026)
Synthesizing Text-to-SQL Data from Weak and Strong LLMs
by: Yang, Jiaxi, et al.
Published: (2024)
by: Yang, Jiaxi, et al.
Published: (2024)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
by: Zheng, Zihao, et al.
Published: (2024)
by: Zheng, Zihao, et al.
Published: (2024)
Debate Helps Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
Selective Weak-to-Strong Generalization
by: Lang, Hao, et al.
Published: (2025)
by: Lang, Hao, et al.
Published: (2025)
How and where does CLIP process negation?
by: Quantmeyer, Vincent, et al.
Published: (2024)
by: Quantmeyer, Vincent, et al.
Published: (2024)
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025)
by: Liao, Chonghua, et al.
Published: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Similar Items
-
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
by: Lang, Hao, et al.
Published: (2023) -
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
by: Cai, Yucheng, et al.
Published: (2024) -
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025) -
Iterative Forward Tuning Boosts In-Context Learning in Language Models
by: Yang, Jiaxi, et al.
Published: (2023) -
Multi-dimensional Evaluation of Empathetic Dialog Responses
by: Xu, Zhichao, et al.
Published: (2024)