DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Zhichao, Hui, Binyuan, Yang, Min, Huang, Fei, Li, Yongbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
von: Lang, Hao, et al.
Veröffentlicht: (2023)
von: Lang, Hao, et al.
Veröffentlicht: (2023)
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
von: Cai, Yucheng, et al.
Veröffentlicht: (2024)
von: Cai, Yucheng, et al.
Veröffentlicht: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Iterative Forward Tuning Boosts In-Context Learning in Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023)
Multi-dimensional Evaluation of Empathetic Dialog Responses
von: Xu, Zhichao, et al.
Veröffentlicht: (2024)
von: Xu, Zhichao, et al.
Veröffentlicht: (2024)
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
von: Rabbani, Parisa, et al.
Veröffentlicht: (2026)
von: Rabbani, Parisa, et al.
Veröffentlicht: (2026)
A Preliminary Study of the Intrinsic Relationship between Complexity and Alignment
von: Zhao, Yingxiu, et al.
Veröffentlicht: (2023)
von: Zhao, Yingxiu, et al.
Veröffentlicht: (2023)
Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
von: Chew, Oscar, et al.
Veröffentlicht: (2026)
MobileCLIP2: Improving Multi-Modal Reinforced Training
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
von: Faghri, Fartash, et al.
Veröffentlicht: (2025)
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
von: Asgarov, Ali, et al.
Veröffentlicht: (2024)
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
von: Xie, Jingyou, et al.
Veröffentlicht: (2024)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
von: Csizmadia, Daniel, et al.
Veröffentlicht: (2025)
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
von: Chen, Junjie, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
Demystifying CLIP Data
von: Xu, Hu, et al.
Veröffentlicht: (2023)
von: Xu, Hu, et al.
Veröffentlicht: (2023)
Jina CLIP: Your CLIP Model Is Also Your Text Retriever
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
von: Koukounas, Andreas, et al.
Veröffentlicht: (2024)
LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
von: Huang, Weiquan, et al.
Veröffentlicht: (2024)
TiC-CLIP: Continual Training of CLIP Models
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
von: Garg, Saurabh, et al.
Veröffentlicht: (2023)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
von: Wang, Enguang, et al.
Veröffentlicht: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
von: Zhang, Xinghua, et al.
Veröffentlicht: (2024)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
von: Cao, Anh-Quan, et al.
Veröffentlicht: (2024)
Meta CLIP 2: A Worldwide Scaling Recipe
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
One-Shot Learning as Instruction Data Prospector for Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2023)
von: Li, Yunshui, et al.
Veröffentlicht: (2023)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
Entriever: Energy-based Retriever for Knowledge-Grounded Dialog Systems
von: Cai, Yucheng, et al.
Veröffentlicht: (2025)
von: Cai, Yucheng, et al.
Veröffentlicht: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
Language-Conditioned Visual Grounding with CLIP Multilingual
von: de Curtò, J., et al.
Veröffentlicht: (2026)
von: de Curtò, J., et al.
Veröffentlicht: (2026)
Synthesizing Text-to-SQL Data from Weak and Strong LLMs
von: Yang, Jiaxi, et al.
Veröffentlicht: (2024)
von: Yang, Jiaxi, et al.
Veröffentlicht: (2024)
KNOWCOMP POKEMON Team at DialAM-2024: A Two-Stage Pipeline for Detecting Relations in Dialogical Argument Mining
von: Zheng, Zihao, et al.
Veröffentlicht: (2024)
von: Zheng, Zihao, et al.
Veröffentlicht: (2024)
Debate Helps Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
Selective Weak-to-Strong Generalization
von: Lang, Hao, et al.
Veröffentlicht: (2025)
von: Lang, Hao, et al.
Veröffentlicht: (2025)
How and where does CLIP process negation?
von: Quantmeyer, Vincent, et al.
Veröffentlicht: (2024)
von: Quantmeyer, Vincent, et al.
Veröffentlicht: (2024)
MOA: Multi-Objective Alignment for Role-Playing Agents
von: Liao, Chonghua, et al.
Veröffentlicht: (2025)
von: Liao, Chonghua, et al.
Veröffentlicht: (2025)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
von: Lang, Hao, et al.
Veröffentlicht: (2023) -
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG)
von: Cai, Yucheng, et al.
Veröffentlicht: (2024) -
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025) -
Iterative Forward Tuning Boosts In-Context Learning in Language Models
von: Yang, Jiaxi, et al.
Veröffentlicht: (2023) -
Multi-dimensional Evaluation of Empathetic Dialog Responses
von: Xu, Zhichao, et al.
Veröffentlicht: (2024)