Vision language models are blind: Failing to translate detailed visual features into words
Fuente:
arXiv
Saved in:
| Main Authors: | Rahmanzadehgervi, Pooyan, Bolton, Logan, Taesiri, Mohammad Reza, Nguyen, Anh Totti |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026)
by: Collins, Brandon, et al.
Published: (2026)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025)
by: Nguyen, Hung Huy, et al.
Published: (2025)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024)
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026)
by: Pooyan R, et al.
Published: (2026)
GlitchBench: Can large multimodal models detect video game glitches?
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
by: Taesiri, Mohammad Reza, et al.
Published: (2023)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
Vision Language Models are Biased
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Allowing humans to interactively guide machines where to look does not always improve human-AI team's classification accuracy
by: Nguyen, Giang, et al.
Published: (2024)
by: Nguyen, Giang, et al.
Published: (2024)
Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma
by: Pastor-Naranjo, Alvaro, et al.
Published: (2025)
by: Pastor-Naranjo, Alvaro, et al.
Published: (2025)
Vision language models are unreliable at trivial spatial cognition
by: Khemlani, Sangeet, et al.
Published: (2025)
by: Khemlani, Sangeet, et al.
Published: (2025)
RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness
by: Park, Junghyun, et al.
Published: (2025)
by: Park, Junghyun, et al.
Published: (2025)
Vision language models have difficulty recognizing virtual objects
by: Tran, Tyler, et al.
Published: (2025)
by: Tran, Tyler, et al.
Published: (2025)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
WAVER: Writing-style Agnostic Text-Video Retrieval via Distilling Vision-Language Models Through Open-Vocabulary Knowledge
by: Le, Huy, et al.
Published: (2023)
by: Le, Huy, et al.
Published: (2023)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine
by: Wu, Yuan, et al.
Published: (2026)
by: Wu, Yuan, et al.
Published: (2026)
Leveraging Habitat Information for Fine-grained Bird Identification
by: Nguyen, Tin, et al.
Published: (2023)
by: Nguyen, Tin, et al.
Published: (2023)
Assessing Color Vision Test in Large Vision-language Models
by: Ye, Hongfei, et al.
Published: (2025)
by: Ye, Hongfei, et al.
Published: (2025)
Anti-I2V: Safeguarding your photos from malicious image-to-video generation
by: Vu, Duc, et al.
Published: (2026)
by: Vu, Duc, et al.
Published: (2026)
FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping
by: Fateh, Amirreza, et al.
Published: (2024)
by: Fateh, Amirreza, et al.
Published: (2024)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
by: Nguyen-Nhu, Tinh-Anh, et al.
Published: (2025)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
by: Nguyen-Truong, Hai, et al.
Published: (2024)
by: Nguyen-Truong, Hai, et al.
Published: (2024)
UNBOX: Unveiling Black-box visual models with Natural-language
by: Carnemolla, Simone, et al.
Published: (2026)
by: Carnemolla, Simone, et al.
Published: (2026)
VideoGameBunny: Towards vision assistants for video games
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
by: Taesiri, Mohammad Reza, et al.
Published: (2024)
UlcerGPT: A Multimodal Approach Leveraging Large Language and Vision Models for Diabetic Foot Ulcer Image Transcription
by: Basiri, Reza, et al.
Published: (2024)
by: Basiri, Reza, et al.
Published: (2024)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
by: Meseguer, Pablo, et al.
Published: (2024)
by: Meseguer, Pablo, et al.
Published: (2024)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
by: QI, Anbin, et al.
Published: (2024)
by: QI, Anbin, et al.
Published: (2024)
Evaluating Large Vision-language Models for Surgical Tool Detection
by: Poudel, Nakul, et al.
Published: (2026)
by: Poudel, Nakul, et al.
Published: (2026)
B-score: Detecting biases in large language models using response history
by: Vo, An, et al.
Published: (2025)
by: Vo, An, et al.
Published: (2025)
CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models
by: Nguyen, Quang-Binh, et al.
Published: (2025)
by: Nguyen, Quang-Binh, et al.
Published: (2025)
MedSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
by: Pham, Trong-Thang, et al.
Published: (2026)
by: Pham, Trong-Thang, et al.
Published: (2026)
HoT: Highlighted Chain of Thought for Referencing Supporting Facts from Inputs
by: Nguyen, Tin, et al.
Published: (2025)
by: Nguyen, Tin, et al.
Published: (2025)
Similar Items
-
SketchVLM: Vision language models can annotate images to explain thoughts and guide users
by: Collins, Brandon, et al.
Published: (2026) -
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
by: Nguyen, Hung Huy, et al.
Published: (2025) -
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025) -
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
by: Rahmanzadehgervi, Pooyan, et al.
Published: (2024) -
anguyen8/vision-llms-are-blind: official
by: Pooyan R, et al.
Published: (2026)