Are VLMs Really Blind

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Ayush, Gupta, Mansi, Garg, Shivank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909370511523840
author Singh, Ayush
Gupta, Mansi
Garg, Shivank
author_facet Singh, Ayush
Gupta, Mansi
Garg, Shivank
contents Vision Language Models excel in handling a wide range of complex tasks, including Optical Character Recognition (OCR), Visual Question Answering (VQA), and advanced geometric reasoning. However, these models fail to perform well on low-level basic visual tasks which are especially easy for humans. Our goal in this work was to determine if these models are truly "blind" to geometric reasoning or if there are ways to enhance their capabilities in this area. Our work presents a novel automatic pipeline designed to extract key information from images in response to specific questions. Instead of just relying on direct VQA, we use question-derived keywords to create a caption that highlights important details in the image related to the question. This caption is then used by a language model to provide a precise answer to the question without requiring external fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22029
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Are VLMs Really Blind
Singh, Ayush
Gupta, Mansi
Garg, Shivank
Computation and Language
Computer Vision and Pattern Recognition
Vision Language Models excel in handling a wide range of complex tasks, including Optical Character Recognition (OCR), Visual Question Answering (VQA), and advanced geometric reasoning. However, these models fail to perform well on low-level basic visual tasks which are especially easy for humans. Our goal in this work was to determine if these models are truly "blind" to geometric reasoning or if there are ways to enhance their capabilities in this area. Our work presents a novel automatic pipeline designed to extract key information from images in response to specific questions. Instead of just relying on direct VQA, we use question-derived keywords to create a caption that highlights important details in the image related to the question. This caption is then used by a language model to provide a precise answer to the question without requiring external fine-tuning.
title Are VLMs Really Blind
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.22029