Freeze and Reveal: Exposing Modality Bias in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kavuri, Vivek Hruday, Karanam, Vysishtya, Venkamsetty, Venkata Jahnavi, Madumadukala, Kriti, Darur, Lakshmipathi Balaji, Kumaraguru, Ponnurangam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026)
by: Darur, Balaji, et al.
Published: (2026)
AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
by: Szu-Tu, Li-Zhong, et al.
Published: (2025)
by: Szu-Tu, Li-Zhong, et al.
Published: (2025)
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
by: Sikarwar, Ankur, et al.
Published: (2026)
by: Sikarwar, Ankur, et al.
Published: (2026)
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
by: Shukla, Tripti, et al.
Published: (2024)
by: Shukla, Tripti, et al.
Published: (2024)
Think$^{2}$: Grounded Metacognitive Reasoning in Large Language Models
by: Elenjical, Abraham Paul, et al.
Published: (2026)
by: Elenjical, Abraham Paul, et al.
Published: (2026)
SPIRIT: Short-term Prediction of solar IRradIance for zero-shot Transfer learning using Foundation Models
by: Mishra, Aditya, et al.
Published: (2025)
by: Mishra, Aditya, et al.
Published: (2025)
Training-free Color-Style Disentanglement for Constrained Text-to-Image Synthesis
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles
by: Bugaud, Zacharie
Published: (2026)
by: Bugaud, Zacharie
Published: (2026)
NEAT: Concept driven Neuron Attribution in LLMs
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
by: Garg, Rahul, et al.
Published: (2024)
by: Garg, Rahul, et al.
Published: (2024)
LABELING COPILOT: A Deep Research Agent for Automated Data Curation in Computer Vision
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
Random Representations Outperform Online Continually Learned Representations
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
by: Kwon, JuneHyoung, et al.
Published: (2025)
by: Kwon, JuneHyoung, et al.
Published: (2025)
KTCR: Improving Implicit Hate Detection with Knowledge Transfer driven Concept Refinement
by: Garg, Samarth, et al.
Published: (2024)
by: Garg, Samarth, et al.
Published: (2024)
TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Pathological Truth Bias in Vision-Language Models
by: Thube, Yash
Published: (2025)
by: Thube, Yash
Published: (2025)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026)
by: Paruchuri, Akshay, et al.
Published: (2026)
LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
Egocentric Bias in Vision-Language Models
by: Wang, Maijunxian, et al.
Published: (2026)
by: Wang, Maijunxian, et al.
Published: (2026)
Few Shot Class Incremental Learning using Vision-Language models
by: Kumar, Anurag, et al.
Published: (2024)
by: Kumar, Anurag, et al.
Published: (2024)
The Bias of Harmful Label Associations in Vision-Language Models
by: Hazirbas, Caner, et al.
Published: (2024)
by: Hazirbas, Caner, et al.
Published: (2024)
Stable Diffusion Exposed: Gender Bias from Prompt to Image
by: Wu, Yankun, et al.
Published: (2023)
by: Wu, Yankun, et al.
Published: (2023)
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
by: Schrodi, Simon, et al.
Published: (2024)
by: Schrodi, Simon, et al.
Published: (2024)
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
by: Liu, Hanqing, et al.
Published: (2025)
by: Liu, Hanqing, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
MorphoFlow: Sparse-Supervised Generative Shape Modeling with Adaptive Latent Relevance
by: Karanam, Mokshagna Sai Teja, et al.
Published: (2026)
by: Karanam, Mokshagna Sai Teja, et al.
Published: (2026)
Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
by: Agarwal, Aishwarya, et al.
Published: (2025)
by: Agarwal, Aishwarya, et al.
Published: (2025)
Learning 3D Texture-Aware Representations for Parsing Diverse Human Clothing and Body Parts
by: Chhatre, Kiran, et al.
Published: (2025)
by: Chhatre, Kiran, et al.
Published: (2025)
LiteEmbed: Adapting CLIP to Rare Classes
by: Agarwal, Aishwarya, et al.
Published: (2026)
by: Agarwal, Aishwarya, et al.
Published: (2026)
Mesh2SSM++: A Probabilistic Framework for Unsupervised Learning of Statistical Shape Model of Anatomies from Surface Meshes
by: Iyer, Krithika, et al.
Published: (2025)
by: Iyer, Krithika, et al.
Published: (2025)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
by: Imran, Muhammad, et al.
Published: (2025)
by: Imran, Muhammad, et al.
Published: (2025)
Cross-Modal Attention Guided Unlearning in Vision-Language Models
by: Bhaila, Karuna, et al.
Published: (2025)
by: Bhaila, Karuna, et al.
Published: (2025)
Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
by: Atabuzzaman, Md., et al.
Published: (2025)
by: Atabuzzaman, Md., et al.
Published: (2025)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
VLM-UQBench: A Benchmark for Modality-Specific and Cross-Modality Uncertainties in Vision Language Models
by: Wang, Chenyu, et al.
Published: (2026)
by: Wang, Chenyu, et al.
Published: (2026)
Similar Items
-
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
by: Darur, Balaji, et al.
Published: (2026) -
AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models
by: Agarwal, Aishwarya, et al.
Published: (2024) -
Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
by: Szu-Tu, Li-Zhong, et al.
Published: (2025) -
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
by: Sikarwar, Ankur, et al.
Published: (2026) -
Test-time Conditional Text-to-Image Synthesis Using Diffusion Models
by: Shukla, Tripti, et al.
Published: (2024)