What to align in multimodal contrastive learning?
Fuente:
arXiv
Saved in:
| Main Authors: | Dufumier, Benoit, Castillo-Navarro, Javiera, Tuia, Devis, Thiran, Jean-Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Cross-Modal Learning of Housing Quality in Amsterdam
by: Levering, Alex, et al.
Published: (2024)
by: Levering, Alex, et al.
Published: (2024)
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Questions beyond Pixels: Integrating Commonsense Knowledge in Visual Question Generation for Remote Sensing
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
SepVAE: a contrastive VAE to separate pathological patterns from healthy ones
by: Louiset, Robin, et al.
Published: (2023)
by: Louiset, Robin, et al.
Published: (2023)
Uncertainty modeling for fine-tuned implicit functions
by: Susmelj, Anna, et al.
Published: (2024)
by: Susmelj, Anna, et al.
Published: (2024)
Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
by: Porta, Hugo, et al.
Published: (2024)
by: Porta, Hugo, et al.
Published: (2024)
Contrastive-to-Self-Supervised: A Two-Stage Framework for Script Similarity Learning
by: Roman, Claire, et al.
Published: (2026)
by: Roman, Claire, et al.
Published: (2026)
Review of multimodal machine learning approaches in healthcare
by: Krones, Felix, et al.
Published: (2024)
by: Krones, Felix, et al.
Published: (2024)
What's Holding Back Latent Visual Reasoning?
by: Viveiros, André G., et al.
Published: (2026)
by: Viveiros, André G., et al.
Published: (2026)
EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
by: Zermatten, Valerie, et al.
Published: (2025)
by: Zermatten, Valerie, et al.
Published: (2025)
CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
by: Chambon, Pierre, et al.
Published: (2024)
by: Chambon, Pierre, et al.
Published: (2024)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
What Makes a Maze Look Like a Maze?
by: Hsu, Joy, et al.
Published: (2024)
by: Hsu, Joy, et al.
Published: (2024)
Reproducible scaling laws for contrastive language-image learning
by: Cherti, Mehdi, et al.
Published: (2022)
by: Cherti, Mehdi, et al.
Published: (2022)
SpinQuant: LLM quantization with learned rotations
by: Liu, Zechun, et al.
Published: (2024)
by: Liu, Zechun, et al.
Published: (2024)
TRUST: Leveraging Text Robustness for Unsupervised Domain Adaptation
by: Litrico, Mattia, et al.
Published: (2025)
by: Litrico, Mattia, et al.
Published: (2025)
Human-like object concept representations emerge naturally in multimodal large language models
by: Du, Changde, et al.
Published: (2024)
by: Du, Changde, et al.
Published: (2024)
HEMM: Holistic Evaluation of Multimodal Foundation Models
by: Liang, Paul Pu, et al.
Published: (2024)
by: Liang, Paul Pu, et al.
Published: (2024)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
Towards Privacy-Aware Sign Language Translation at Scale
by: Rust, Phillip, et al.
Published: (2024)
by: Rust, Phillip, et al.
Published: (2024)
IoT-LM: Large Multisensory Language Models for the Internet of Things
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
MultiIoT: Benchmarking Machine Learning for the Internet of Things
by: Mo, Shentong, et al.
Published: (2023)
by: Mo, Shentong, et al.
Published: (2023)
Do multimodal models imagine electric sheep?
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
by: Ramakrishnan, Santhosh Kumar, et al.
Published: (2026)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
by: Englebert, Alexandre, et al.
Published: (2024)
by: Englebert, Alexandre, et al.
Published: (2024)
Automatic Discovery of Disease Subgroups by Contrasting with Healthy Controls
by: Louiset, Robin, et al.
Published: (2026)
by: Louiset, Robin, et al.
Published: (2026)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
CFM: Language-aligned Concept Foundation Model for Vision
by: Wittenmayer, Kai, et al.
Published: (2026)
by: Wittenmayer, Kai, et al.
Published: (2026)
VaPR -- Vision-language Preference alignment for Reasoning
by: Wadhawan, Rohan, et al.
Published: (2025)
by: Wadhawan, Rohan, et al.
Published: (2025)
Anatomical Foundation Models for Brain MRIs
by: Barbano, Carlo Alberto, et al.
Published: (2024)
by: Barbano, Carlo Alberto, et al.
Published: (2024)
Capabilities of Gemini Models in Medicine
by: Saab, Khaled, et al.
Published: (2024)
by: Saab, Khaled, et al.
Published: (2024)
NLLG Quarterly arXiv Report 09/24: What are the most influential current AI Papers?
by: Leiter, Christoph, et al.
Published: (2024)
by: Leiter, Christoph, et al.
Published: (2024)
From Classification to Segmentation with Explainable AI: A Study on Crack Detection and Growth Monitoring
by: Forest, Florent, et al.
Published: (2023)
by: Forest, Florent, et al.
Published: (2023)
Can multimodal representation learning by alignment preserve modality-specific information?
by: Thoreau, Romain, et al.
Published: (2025)
by: Thoreau, Romain, et al.
Published: (2025)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
by: Ma, Jingkun, et al.
Published: (2024)
by: Ma, Jingkun, et al.
Published: (2024)
Similar Items
-
ConVQG: Contrastive Visual Question Generation with Multimodal Guidance
by: Mi, Li, et al.
Published: (2024) -
Cross-Modal Learning of Housing Quality in Amsterdam
by: Levering, Alex, et al.
Published: (2024) -
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
by: Oota, Subba Reddy, et al.
Published: (2025) -
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024) -
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)