VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
Fuente:
arXiv
Saved in:
| Main Authors: | Shahgir, Haz Sameen, Chen, Xiaofu, Fu, Yu, Shayegani, Erfan, Abu-Ghazaleh, Nael, Kementchedjhieva, Yova, Dong, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025)
by: Shahariar, G M, et al.
Published: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
by: Salazar, Israfel, et al.
Published: (2025)
by: Salazar, Israfel, et al.
Published: (2025)
Harnessing the Unseen: The Hidden Influence of Intrinsic Knowledge in Long-Context Language Models
by: Fu, Yu, et al.
Published: (2025)
by: Fu, Yu, et al.
Published: (2025)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
by: Bachu, Saketh, et al.
Published: (2024)
by: Bachu, Saketh, et al.
Published: (2024)
ExpertGenQA: Open-ended QA generation in Specialized Domains
by: Shahgir, Haz Sameen, et al.
Published: (2025)
by: Shahgir, Haz Sameen, et al.
Published: (2025)
Do Reasoning LLMs Refuse What They Infer in Long Contexts?
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025)
by: Shoer, Belal, et al.
Published: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024)
by: Huzaifa, Muhammad, et al.
Published: (2024)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023)
by: Li, Jiaang, et al.
Published: (2023)
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
by: Ibrahim, George, et al.
Published: (2025)
by: Ibrahim, George, et al.
Published: (2025)
Answerability in Retrieval-Augmented Open-Domain Question Answering
by: Abdumalikov, Rustam, et al.
Published: (2024)
by: Abdumalikov, Rustam, et al.
Published: (2024)
LLMs Can Compensate for Deficiencies in Visual Representations
by: Takishita, Sho, et al.
Published: (2025)
by: Takishita, Sho, et al.
Published: (2025)
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
by: Shahgir, Haz Sameen, et al.
Published: (2023)
by: Shahgir, Haz Sameen, et al.
Published: (2023)
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs
by: Mahfuz, Tamzeed, et al.
Published: (2024)
by: Mahfuz, Tamzeed, et al.
Published: (2024)
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023)
by: Slocum, Carter, et al.
Published: (2023)
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
Connecting the Dots: Leveraging Spatio-Temporal Graph Neural Networks for Accurate Bangla Sign Language Recognition
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
by: Geng, Jiahui, et al.
Published: (2024)
by: Geng, Jiahui, et al.
Published: (2024)
A Discriminative Latent-Variable Model for Bilingual Lexicon Induction
by: Ruder, Sebastian, et al.
Published: (2018)
by: Ruder, Sebastian, et al.
Published: (2018)
Overcoming Vocabulary Constraints with Pixel-level Fallback
by: Lotz, Jonas F., et al.
Published: (2025)
by: Lotz, Jonas F., et al.
Published: (2025)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
A Real-Time System to Populate FRA Form 57 from News
by: Lim, Chansong, et al.
Published: (2025)
by: Lim, Chansong, et al.
Published: (2025)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
by: Imam, Raza, et al.
Published: (2025)
by: Imam, Raza, et al.
Published: (2025)
MuLan: A Study of Fact Mutability in Language Models
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Co(ve)rtex: ML Models as storage channels and their (mis-)applications
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
Leveraging Complementary Attention maps in vision transformers for OCT image analysis
by: Shahgir, Haz Sameen, et al.
Published: (2023)
by: Shahgir, Haz Sameen, et al.
Published: (2023)
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
JEEM: Vision-Language Understanding in Four Arabic Dialects
by: Kadaoui, Karima, et al.
Published: (2025)
by: Kadaoui, Karima, et al.
Published: (2025)
Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
by: Shi, Youxu, et al.
Published: (2025)
by: Shi, Youxu, et al.
Published: (2025)
AbFlowNet: Optimizing Antibody-Antigen Binding Energy via Diffusion-GFlowNet Fusion
by: Abir, Abrar Rahman, et al.
Published: (2025)
by: Abir, Abrar Rahman, et al.
Published: (2025)
Similar Items
-
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025) -
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025) -
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025) -
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025) -
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
by: Salazar, Israfel, et al.
Published: (2025)