Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jialou, Zhu, Manli, Li, Yulei, Li, Honglei, Yang, Longzhi, Woo, Wai Lok
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911821566312448
author Wang, Jialou
Zhu, Manli
Li, Yulei
Li, Honglei
Yang, Longzhi
Woo, Wai Lok
author_facet Wang, Jialou
Zhu, Manli
Li, Yulei
Li, Honglei
Yang, Longzhi
Woo, Wai Lok
contents Localization plays a crucial role in enhancing the practicality and precision of VQA systems. By enabling fine-grained identification and interaction with specific parts of an object, it significantly improves the system's ability to provide contextually relevant and spatially accurate responses, crucial for applications in dynamic environments like robotics and augmented reality. However, traditional systems face challenges in accurately mapping objects within images to generate nuanced and spatially aware responses. In this work, we introduce "Detect2Interact", which addresses these challenges by introducing an advanced approach for fine-grained object visual key field detection. First, we use the segment anything model (SAM) to generate detailed spatial maps of objects in images. Next, we use Vision Studio to extract semantic object descriptions. Third, we employ GPT-4's common sense knowledge, bridging the gap between an object's semantics and its spatial map. As a result, Detect2Interact achieves consistent qualitative results on object key field detection across extensive test cases and outperforms the existing VQA system with object detection by providing a more reasonable and finer visual representation.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01151
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
Wang, Jialou
Zhu, Manli
Li, Yulei
Li, Honglei
Yang, Longzhi
Woo, Wai Lok
Computer Vision and Pattern Recognition
Localization plays a crucial role in enhancing the practicality and precision of VQA systems. By enabling fine-grained identification and interaction with specific parts of an object, it significantly improves the system's ability to provide contextually relevant and spatially accurate responses, crucial for applications in dynamic environments like robotics and augmented reality. However, traditional systems face challenges in accurately mapping objects within images to generate nuanced and spatially aware responses. In this work, we introduce "Detect2Interact", which addresses these challenges by introducing an advanced approach for fine-grained object visual key field detection. First, we use the segment anything model (SAM) to generate detailed spatial maps of objects in images. Next, we use Vision Studio to extract semantic object descriptions. Third, we employ GPT-4's common sense knowledge, bridging the gap between an object's semantics and its spatial map. As a result, Detect2Interact achieves consistent qualitative results on object key field detection across extensive test cases and outperforms the existing VQA system with object detection by providing a more reasonable and finer visual representation.
title Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.01151