Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Li, Yang, Diji, Zhong, Sijia, Tholeti, Kalyana Suma Sree, Ding, Lei, Zhang, Yi, Gilpin, Leilani H.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2411.00394
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910680549949440
author Liu, Li
Yang, Diji
Zhong, Sijia
Tholeti, Kalyana Suma Sree
Ding, Lei
Zhang, Yi
Gilpin, Leilani H.
author_facet Liu, Li
Yang, Diji
Zhong, Sijia
Tholeti, Kalyana Suma Sree
Ding, Lei
Zhang, Yi
Gilpin, Leilani H.
contents In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the sufficiency of the information. To investigate this gap, we identify a critical and challenging task in the Visual Question Answering (VQA) scenario: can VLMs indicate how to adjust an image when the visual information is insufficient to answer a question? This capability is especially valuable for assisting visually impaired individuals who often need guidance to capture images correctly. To evaluate this capability of current VLMs, we introduce a human-labeled dataset as a benchmark for this task. Additionally, we present an automated framework that generates synthetic training data by simulating ``where to know'' scenarios. Our empirical results show significant performance improvements in mainstream VLMs when fine-tuned with this synthetic data. This study demonstrates the potential to narrow the gap between information assessment and acquisition in VLMs, bringing their performance closer to humans.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00394
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Right this way: Can VLMs Guide Us to See More to Answer Questions?
Liu, Li
Yang, Diji
Zhong, Sijia
Tholeti, Kalyana Suma Sree
Ding, Lei
Zhang, Yi
Gilpin, Leilani H.
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
In question-answering scenarios, humans can assess whether the available information is sufficient and seek additional information if necessary, rather than providing a forced answer. In contrast, Vision Language Models (VLMs) typically generate direct, one-shot responses without evaluating the sufficiency of the information. To investigate this gap, we identify a critical and challenging task in the Visual Question Answering (VQA) scenario: can VLMs indicate how to adjust an image when the visual information is insufficient to answer a question? This capability is especially valuable for assisting visually impaired individuals who often need guidance to capture images correctly. To evaluate this capability of current VLMs, we introduce a human-labeled dataset as a benchmark for this task. Additionally, we present an automated framework that generates synthetic training data by simulating ``where to know'' scenarios. Our empirical results show significant performance improvements in mainstream VLMs when fine-tuned with this synthetic data. This study demonstrates the potential to narrow the gap between information assessment and acquisition in VLMs, bringing their performance closer to humans.
title Right this way: Can VLMs Guide Us to See More to Answer Questions?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.00394