VISREAS: Complex Visual Reasoning with Unanswerable Questions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Akter, Syeda Nahida, Lee, Sangwu, Chang, Yingshan, Bisk, Yonatan, Nyberg, Eric
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914715827961856
author Akter, Syeda Nahida
Lee, Sangwu
Chang, Yingshan
Bisk, Yonatan
Nyberg, Eric
author_facet Akter, Syeda Nahida
Lee, Sangwu
Chang, Yingshan
Bisk, Yonatan
Nyberg, Eric
contents Verifying a question's validity before answering is crucial in real-world applications, where users may provide imperfect instructions. In this scenario, an ideal model should address the discrepancies in the query and convey them to the users rather than generating the best possible answer. Addressing this requirement, we introduce a new compositional visual question-answering dataset, VISREAS, that consists of answerable and unanswerable visual queries formulated by traversing and perturbing commonalities and differences among objects, attributes, and relations. VISREAS contains 2.07M semantically diverse queries generated automatically using Visual Genome scene graphs. The unique feature of this task, validating question answerability with respect to an image before answering, and the poor performance of state-of-the-art models inspired the design of a new modular baseline, LOGIC2VISION that reasons by producing and executing pseudocode without any external modules to generate the answer. LOGIC2VISION outperforms generative models in VISREAS (+4.82% over LLaVA-1.5; +12.23% over InstructBLIP) and achieves a significant gain in performance against the classification models.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10534
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VISREAS: Complex Visual Reasoning with Unanswerable Questions
Akter, Syeda Nahida
Lee, Sangwu
Chang, Yingshan
Bisk, Yonatan
Nyberg, Eric
Computer Vision and Pattern Recognition
Artificial Intelligence
Verifying a question's validity before answering is crucial in real-world applications, where users may provide imperfect instructions. In this scenario, an ideal model should address the discrepancies in the query and convey them to the users rather than generating the best possible answer. Addressing this requirement, we introduce a new compositional visual question-answering dataset, VISREAS, that consists of answerable and unanswerable visual queries formulated by traversing and perturbing commonalities and differences among objects, attributes, and relations. VISREAS contains 2.07M semantically diverse queries generated automatically using Visual Genome scene graphs. The unique feature of this task, validating question answerability with respect to an image before answering, and the poor performance of state-of-the-art models inspired the design of a new modular baseline, LOGIC2VISION that reasons by producing and executing pseudocode without any external modules to generate the answer. LOGIC2VISION outperforms generative models in VISREAS (+4.82% over LLaVA-1.5; +12.23% over InstructBLIP) and achieves a significant gain in performance against the classification models.
title VISREAS: Complex Visual Reasoning with Unanswerable Questions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2403.10534