ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Ruiping, Zhang, Jiaming, Schön, Angela, Müller, Karin, Zheng, Junwei, Yang, Kailun, Guo, Anhong, Gerling, Kathrin, Stiefelhagen, Rainer
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916713804595200
author Liu, Ruiping
Zhang, Jiaming
Schön, Angela
Müller, Karin
Zheng, Junwei
Yang, Kailun
Guo, Anhong
Gerling, Kathrin
Stiefelhagen, Rainer
author_facet Liu, Ruiping
Zhang, Jiaming
Schön, Angela
Müller, Karin
Zheng, Junwei
Yang, Kailun
Guo, Anhong
Gerling, Kathrin
Stiefelhagen, Rainer
contents Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03118
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
Liu, Ruiping
Zhang, Jiaming
Schön, Angela
Müller, Karin
Zheng, Junwei
Yang, Kailun
Guo, Anhong
Gerling, Kathrin
Stiefelhagen, Rainer
Human-Computer Interaction
Computer Vision and Pattern Recognition
Searching for objects in unfamiliar scenarios is a challenging task for blind people. It involves specifying the target object, detecting it, and then gathering detailed information according to the user's intent. However, existing description- and detection-based assistive technologies do not sufficiently support the multifaceted nature of interactive object search tasks. We present ObjectFinder, an open-vocabulary wearable assistive system for interactive object search by blind people. ObjectFinder allows users to query target objects using flexible wording. Once the target object is detected, it provides egocentric localization information in real-time, including distance and direction. Users can then initiate different branches to gather detailed information based on their intent towards the target object, such as navigating to it or perceiving its surroundings. ObjectFinder is powered by a seamless combination of open-vocabulary models, namely an open-vocabulary object detector and a multimodal large language model. The ObjectFinder design concept and its development were carried out in collaboration with a blind co-designer. To evaluate ObjectFinder, we conducted an exploratory user study with eight blind participants. We compared ObjectFinder to BeMyAI and Google Lookout, popular description- and detection-based assistive applications. Our findings indicate that most participants felt more independent with ObjectFinder and preferred it for object search, as it enhanced scene context gathering and navigation, and allowed for active target identification. Finally, we discuss the implications for future assistive systems to support interactive object search.
title ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.03118