Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
Fuente:
arXiv
Salvato in:
| Autori principali: | Eslami, Sedigheh, de Melo, Gerard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MobileCLIP2: Improving Multi-Modal Reinforced Training
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
di: Faghri, Fartash, et al.
Pubblicazione: (2025)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
di: Mistretta, Marco, et al.
Pubblicazione: (2025)
di: Mistretta, Marco, et al.
Pubblicazione: (2025)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
di: Song, Zijia, et al.
Pubblicazione: (2024)
di: Song, Zijia, et al.
Pubblicazione: (2024)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
di: Wei, Lai, et al.
Pubblicazione: (2023)
di: Wei, Lai, et al.
Pubblicazione: (2023)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)
Closing the Modality Gap for Mixed Modality Search
di: Li, Binxu, et al.
Pubblicazione: (2025)
di: Li, Binxu, et al.
Pubblicazione: (2025)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
di: Yaras, Can, et al.
Pubblicazione: (2024)
di: Yaras, Can, et al.
Pubblicazione: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
di: Deng, Ailin, et al.
Pubblicazione: (2024)
di: Deng, Ailin, et al.
Pubblicazione: (2024)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
di: Chang, Kai-Po, et al.
Pubblicazione: (2025)
di: Chang, Kai-Po, et al.
Pubblicazione: (2025)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
di: Wang, Xiyao, et al.
Pubblicazione: (2024)
Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
di: Gautam, Somraj, et al.
Pubblicazione: (2025)
di: Gautam, Somraj, et al.
Pubblicazione: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
di: Kim, Sanghwan, et al.
Pubblicazione: (2024)
Enhancing CLIP Conceptual Embedding through Knowledge Distillation
di: Kao, Kuei-Chun
Pubblicazione: (2024)
di: Kao, Kuei-Chun
Pubblicazione: (2024)
MoDE: CLIP Data Experts via Clustering
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
di: Ma, Jiawei, et al.
Pubblicazione: (2024)
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts
di: An, Bang, et al.
Pubblicazione: (2023)
di: An, Bang, et al.
Pubblicazione: (2023)
EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
di: Lin, Dongyan, et al.
Pubblicazione: (2026)
di: Lin, Dongyan, et al.
Pubblicazione: (2026)
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
di: Mantini, Pranav, et al.
Pubblicazione: (2026)
di: Mantini, Pranav, et al.
Pubblicazione: (2026)
Improving Alignment and Robustness with Circuit Breakers
di: Zou, Andy, et al.
Pubblicazione: (2024)
di: Zou, Andy, et al.
Pubblicazione: (2024)
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery
di: Wang, Enguang, et al.
Pubblicazione: (2024)
di: Wang, Enguang, et al.
Pubblicazione: (2024)
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
di: Huang, Hanxun, et al.
Pubblicazione: (2025)
di: Huang, Hanxun, et al.
Pubblicazione: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
di: Wei, Lai, et al.
Pubblicazione: (2025)
di: Wei, Lai, et al.
Pubblicazione: (2025)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
di: Mo, Wentao, et al.
Pubblicazione: (2024)
di: Mo, Wentao, et al.
Pubblicazione: (2024)
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks
di: Hu, Yuanze, et al.
Pubblicazione: (2025)
di: Hu, Yuanze, et al.
Pubblicazione: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
di: Lai, Zhengfeng, et al.
Pubblicazione: (2023)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
di: Liang, Weixin, et al.
Pubblicazione: (2025)
di: Liang, Weixin, et al.
Pubblicazione: (2025)
StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
di: Guo, Ziyu, et al.
Pubblicazione: (2025)
Self-Supervised Visual Preference Alignment
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
ET tu, CLIP? Addressing Common Object Errors for Unseen Environments
di: Byun, Ye Won, et al.
Pubblicazione: (2024)
di: Byun, Ye Won, et al.
Pubblicazione: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
di: Liu, Yiming, et al.
Pubblicazione: (2025)
di: Liu, Yiming, et al.
Pubblicazione: (2025)
Find The Gap: Knowledge Base Reasoning For Visual Question Answering
di: Barezi, Elham J., et al.
Pubblicazione: (2024)
di: Barezi, Elham J., et al.
Pubblicazione: (2024)
Bridging the Gap Between Multimodal Foundation Models and World Models
di: He, Xuehai
Pubblicazione: (2025)
di: He, Xuehai
Pubblicazione: (2025)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
di: Dalvi, Abhishek, et al.
Pubblicazione: (2026)
di: Dalvi, Abhishek, et al.
Pubblicazione: (2026)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
di: Hua, Zhenglin, et al.
Pubblicazione: (2025)
Preliminary Investigations of a Multi-Faceted Robust and Synergistic Approach in Semiconductor Electron Micrograph Analysis: Integrating Vision Transformers with Large Language and Multimodal Models
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2024)
di: Srinivas, Sakhinana Sagar, et al.
Pubblicazione: (2024)
LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
di: Hou, Yuchen, et al.
Pubblicazione: (2026)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
di: Ma, Yunsheng, et al.
Pubblicazione: (2024)
di: Ma, Yunsheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MobileCLIP2: Improving Multi-Modal Reinforced Training
di: Faghri, Fartash, et al.
Pubblicazione: (2025) -
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
di: Mistretta, Marco, et al.
Pubblicazione: (2025) -
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
di: Song, Zijia, et al.
Pubblicazione: (2024) -
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
di: Wei, Lai, et al.
Pubblicazione: (2023) -
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
di: Zhang, Yuhui, et al.
Pubblicazione: (2024)