Vision-Language Models Can't See the Obvious
Fuente:
arXiv
Guardado en:
| Autores principales: | Dahou, Yasser, Huynh, Ngoc Dung, Le-Khac, Phuc H., Para, Wamiq Reyaz, Singh, Ankit, Narayan, Sanath |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
por: Törtei, Brigitta Malagurski, et al.
Publicado: (2025)
por: Törtei, Brigitta Malagurski, et al.
Publicado: (2025)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
por: Chaybouti, Sofian, et al.
Publicado: (2025)
por: Chaybouti, Sofian, et al.
Publicado: (2025)
Falcon Perception
por: Bevli, Aviraj, et al.
Publicado: (2026)
por: Bevli, Aviraj, et al.
Publicado: (2026)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
por: Maniparambil, Mayug, et al.
Publicado: (2024)
por: Maniparambil, Mayug, et al.
Publicado: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
por: Para, Wamiq Reyaz, et al.
Publicado: (2024)
por: Para, Wamiq Reyaz, et al.
Publicado: (2024)
Do Vision and Language Encoders Represent the World Similarly?
por: Maniparambil, Mayug, et al.
Publicado: (2024)
por: Maniparambil, Mayug, et al.
Publicado: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
por: Upadhyay, Ujjwal, et al.
Publicado: (2025)
ViSpeR: Multilingual Audio-Visual Speech Recognition
por: Narayan, Sanath, et al.
Publicado: (2024)
por: Narayan, Sanath, et al.
Publicado: (2024)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
How Well Can Vision Language Models See Image Details?
por: Gou, Chenhui, et al.
Publicado: (2024)
por: Gou, Chenhui, et al.
Publicado: (2024)
Boosting Medical Vision-Language Pretraining via Momentum Self-Distillation under Limited Computing Resources
por: Pham, Phuc, et al.
Publicado: (2025)
por: Pham, Phuc, et al.
Publicado: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
por: Kamath, Amita, et al.
Publicado: (2026)
por: Kamath, Amita, et al.
Publicado: (2026)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
por: Lee, Heekyung, et al.
Publicado: (2025)
por: Lee, Heekyung, et al.
Publicado: (2025)
In the Era of Prompt Learning with Vision-Language Models
por: Jha, Ankit
Publicado: (2024)
por: Jha, Ankit
Publicado: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
por: Guo, Xuyang, et al.
Publicado: (2025)
por: Guo, Xuyang, et al.
Publicado: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
por: Wang, Wenxuan, et al.
Publicado: (2025)
por: Wang, Wenxuan, et al.
Publicado: (2025)
I Can't Believe It's Not Scene Flow!
por: Khatri, Ishan, et al.
Publicado: (2024)
por: Khatri, Ishan, et al.
Publicado: (2024)
Can Vision-Language Models Solve Visual Math Equations?
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
por: Choudhury, Monjoy Narayan, et al.
Publicado: (2025)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
por: Li, Aaron Branson Cigres, et al.
Publicado: (2026)
por: Li, Aaron Branson Cigres, et al.
Publicado: (2026)
Can Graphs Help Vision SSMs See Better?
por: Parikh, Dhruv, et al.
Publicado: (2026)
por: Parikh, Dhruv, et al.
Publicado: (2026)
Unifying Global and Local Scene Entities Modelling for Precise Action Spotting
por: Tran, Kim Hoang, et al.
Publicado: (2024)
por: Tran, Kim Hoang, et al.
Publicado: (2024)
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
por: Huynh, Ngoc Dung, et al.
Publicado: (2024)
por: Huynh, Ngoc Dung, et al.
Publicado: (2024)
Falcon2-11B Technical Report
por: Malartic, Quentin, et al.
Publicado: (2024)
por: Malartic, Quentin, et al.
Publicado: (2024)
GMAT: Grounded Multi-Agent Clinical Description Generation for Text Encoder in Vision-Language MIL for Whole Slide Image Classification
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
por: Quang, Ngoc Bui Lam, et al.
Publicado: (2025)
Visual question answering: from early developments to recent advances -- a survey
por: Huynh, Ngoc Dung, et al.
Publicado: (2025)
por: Huynh, Ngoc Dung, et al.
Publicado: (2025)
Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language Models
por: Mittal, Himangi, et al.
Publicado: (2024)
por: Mittal, Himangi, et al.
Publicado: (2024)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
por: Talon, Davide, et al.
Publicado: (2025)
por: Talon, Davide, et al.
Publicado: (2025)
Can Vision-Language Models See Squares? Text-Recognition Mediates Spatial Reasoning Across Three Model Families
por: Levental, Yuval
Publicado: (2026)
por: Levental, Yuval
Publicado: (2026)
Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning
por: Liang, Dayong, et al.
Publicado: (2025)
por: Liang, Dayong, et al.
Publicado: (2025)
GenFlow: Interactive Modular System for Image Generation
por: Nguyen, Duc-Hung, et al.
Publicado: (2025)
por: Nguyen, Duc-Hung, et al.
Publicado: (2025)
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
por: Tang, Jingyi, et al.
Publicado: (2026)
por: Tang, Jingyi, et al.
Publicado: (2026)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
por: Liu, Chengxin, et al.
Publicado: (2026)
por: Liu, Chengxin, et al.
Publicado: (2026)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
por: Zhou, Heng, et al.
Publicado: (2026)
por: Zhou, Heng, et al.
Publicado: (2026)
Exploring Primitive Visual Measurement Understanding and the Role of Output Format in Learning in Vision-Language Models
por: Yadav, Ankit, et al.
Publicado: (2025)
por: Yadav, Ankit, et al.
Publicado: (2025)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
por: Shinnick, Zachary, et al.
Publicado: (2025)
por: Shinnick, Zachary, et al.
Publicado: (2025)
Multi-modal Generation via Cross-Modal In-Context Learning
por: Kumar, Amandeep, et al.
Publicado: (2024)
por: Kumar, Amandeep, et al.
Publicado: (2024)
BLINK: Multimodal Large Language Models Can See but Not Perceive
por: Fu, Xingyu, et al.
Publicado: (2024)
por: Fu, Xingyu, et al.
Publicado: (2024)
STER-VLM: Spatio-Temporal With Enhanced Reference Vision-Language Models
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
por: Nguyen-Nhu, Tinh-Anh, et al.
Publicado: (2025)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
por: Hu, Xin, et al.
Publicado: (2026)
por: Hu, Xin, et al.
Publicado: (2026)
Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions
por: Wang, Xuesong, et al.
Publicado: (2026)
por: Wang, Xuesong, et al.
Publicado: (2026)
Ejemplares similares
-
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
por: Törtei, Brigitta Malagurski, et al.
Publicado: (2025) -
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
por: Chaybouti, Sofian, et al.
Publicado: (2025) -
Falcon Perception
por: Bevli, Aviraj, et al.
Publicado: (2026) -
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
por: Maniparambil, Mayug, et al.
Publicado: (2024) -
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
por: Para, Wamiq Reyaz, et al.
Publicado: (2024)