Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Schrodi, Simon, Hoffmann, David T., Argus, Max, Fischer, Volker, Brox, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025)
by: Kempf, Elias, et al.
Published: (2025)
Concept Bottleneck Models Without Predefined Concepts
by: Schrodi, Simon, et al.
Published: (2024)
by: Schrodi, Simon, et al.
Published: (2024)
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023)
by: Hoffmann, David T., et al.
Published: (2023)
Common Data Properties Limit Object-Attribute Binding in CLIP
by: Gurung, Bijay, et al.
Published: (2025)
by: Gurung, Bijay, et al.
Published: (2025)
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)
by: Farid, Karim, et al.
Published: (2025)
DITTO: Demonstration Imitation by Trajectory Transformation
by: Heppert, Nick, et al.
Published: (2024)
by: Heppert, Nick, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
Detect, Classify, Act: Categorizing Industrial Anomalies with Multi-Modal Large Language Models
by: Mokhtar, Sassan, et al.
Published: (2025)
by: Mokhtar, Sassan, et al.
Published: (2025)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
On the Domain Robustness of Contrastive Vision-Language Models
by: Koddenbrock, Mario, et al.
Published: (2025)
by: Koddenbrock, Mario, et al.
Published: (2025)
Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
Bridge the Modality and Capability Gaps in Vision-Language Model Selection
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
A Tale of Two Classes: Adapting Supervised Contrastive Learning to Binary Imbalanced Datasets
by: Mildenberger, David, et al.
Published: (2025)
by: Mildenberger, David, et al.
Published: (2025)
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
by: Fahim, Abrar, et al.
Published: (2024)
by: Fahim, Abrar, et al.
Published: (2024)
Anomaly Detection with Conditioned Denoising Diffusion Models
by: Mousakhan, Arian, et al.
Published: (2023)
by: Mousakhan, Arian, et al.
Published: (2023)
See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
by: Kwon, JuneHyoung, et al.
Published: (2025)
by: Kwon, JuneHyoung, et al.
Published: (2025)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
by: Kim, SiWoo, et al.
Published: (2025)
by: Kim, SiWoo, et al.
Published: (2025)
Extending 6D Object Pose Estimators for Stereo Vision
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
by: Dünkel, Olaf, et al.
Published: (2026)
by: Dünkel, Olaf, et al.
Published: (2026)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
by: Xu, Yige, et al.
Published: (2026)
by: Xu, Yige, et al.
Published: (2026)
Label-Efficient LiDAR Semantic Segmentation with 2D-3D Vision Transformer Adapters
by: Hindel, Julia, et al.
Published: (2025)
by: Hindel, Julia, et al.
Published: (2025)
Unsupervised Open-Vocabulary Object Localization in Videos
by: Fan, Ke, et al.
Published: (2023)
by: Fan, Ke, et al.
Published: (2023)
Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CT
by: Ging, Simon, et al.
Published: (2026)
by: Ging, Simon, et al.
Published: (2026)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization
by: Gaur, Gopalji, et al.
Published: (2025)
by: Gaur, Gopalji, et al.
Published: (2025)
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
by: Szu-Tu, Li-Zhong, et al.
Published: (2025)
by: Szu-Tu, Li-Zhong, et al.
Published: (2025)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment
by: Zavras, Angelos, et al.
Published: (2024)
by: Zavras, Angelos, et al.
Published: (2024)
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
by: Bratulić, Jelena, et al.
Published: (2025)
by: Bratulić, Jelena, et al.
Published: (2025)
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
by: Guesmi, Amira, et al.
Published: (2026)
by: Guesmi, Amira, et al.
Published: (2026)
Pathological Truth Bias in Vision-Language Models
by: Thube, Yash
Published: (2025)
by: Thube, Yash
Published: (2025)
Information Router for Mitigating Modality Dominance in Vision-Language Models
by: Kim, Seulgi, et al.
Published: (2026)
by: Kim, Seulgi, et al.
Published: (2026)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models
by: Hu, Yuanwei, et al.
Published: (2026)
by: Hu, Yuanwei, et al.
Published: (2026)
Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings
by: Dhimoïla, Grégoire, et al.
Published: (2026)
by: Dhimoïla, Grégoire, et al.
Published: (2026)
Similar Items
-
When and How Does CLIP Enable Domain and Compositional Generalization?
by: Kempf, Elias, et al.
Published: (2025) -
Concept Bottleneck Models Without Predefined Concepts
by: Schrodi, Simon, et al.
Published: (2024) -
Eureka-Moments in Transformers: Multi-Step Tasks Reveal Softmax Induced Optimization Problems
by: Hoffmann, David T., et al.
Published: (2023) -
Common Data Properties Limit Object-Attribute Binding in CLIP
by: Gurung, Bijay, et al.
Published: (2025) -
What Drives Compositional Generalization? The Importance of Continuous Training Objectives in Visual Generative Models
by: Farid, Karim, et al.
Published: (2025)