Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merchant, Zain, Anwar, Abrar, Wang, Emily, Chattopadhyay, Souti, Thomason, Jesse
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911952697032704
author Merchant, Zain
Anwar, Abrar
Wang, Emily
Chattopadhyay, Souti
Thomason, Jesse
author_facet Merchant, Zain
Anwar, Abrar
Wang, Emily
Chattopadhyay, Souti
Thomason, Jesse
contents Navigating unfamiliar environments presents significant challenges for blind and low-vision (BLV) individuals. In this work, we construct a dataset of images and goals across different scenarios such as searching through kitchens or navigating outdoors. We then investigate how grounded instruction generation methods can provide contextually-relevant navigational guidance to users in these instances. Through a sighted user study, we demonstrate that large pretrained language models can produce correct and useful instructions perceived as beneficial for BLV users. We also conduct a survey and interview with 4 BLV users and observe useful insights on preferences for different instructions based on the scenario.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08219
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
Merchant, Zain
Anwar, Abrar
Wang, Emily
Chattopadhyay, Souti
Thomason, Jesse
Computation and Language
Human-Computer Interaction
Navigating unfamiliar environments presents significant challenges for blind and low-vision (BLV) individuals. In this work, we construct a dataset of images and goals across different scenarios such as searching through kitchens or navigating outdoors. We then investigate how grounded instruction generation methods can provide contextually-relevant navigational guidance to users in these instances. Through a sighted user study, we demonstrate that large pretrained language models can produce correct and useful instructions perceived as beneficial for BLV users. We also conduct a survey and interview with 4 BLV users and observe useful insights on preferences for different instructions based on the scenario.
title Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2407.08219