Enhancing the Travel Experience for People with Visual Impairments through Multimodal Interaction: NaviGPT, A Real-Time AI-Driven Mobile Navigation System

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, He, Falletta, Nicholas J., Xie, Jingyi, Yu, Rui, Lee, Sooyeon, Billah, Syed Masum, Carroll, John M.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910634685235200
author Zhang, He
Falletta, Nicholas J.
Xie, Jingyi
Yu, Rui
Lee, Sooyeon
Billah, Syed Masum
Carroll, John M.
author_facet Zhang, He
Falletta, Nicholas J.
Xie, Jingyi
Yu, Rui
Lee, Sooyeon
Billah, Syed Masum
Carroll, John M.
contents Assistive technologies for people with visual impairments (PVI) have made significant advancements, particularly with the integration of artificial intelligence (AI) and real-time sensor technologies. However, current solutions often require PVI to switch between multiple apps and tools for tasks like image recognition, navigation, and obstacle detection, which can hinder a seamless and efficient user experience. In this paper, we present NaviGPT, a high-fidelity prototype that integrates LiDAR-based obstacle detection, vibration feedback, and large language model (LLM) responses to provide a comprehensive and real-time navigation aid for PVI. Unlike existing applications such as Be My AI and Seeing AI, NaviGPT combines image recognition and contextual navigation guidance into a single system, offering continuous feedback on the user's surroundings without the need for app-switching. Meanwhile, NaviGPT compensates for the response delays of LLM by using location and sensor data, aiming to provide practical and efficient navigation support for PVI in dynamic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04005
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing the Travel Experience for People with Visual Impairments through Multimodal Interaction: NaviGPT, A Real-Time AI-Driven Mobile Navigation System
Zhang, He
Falletta, Nicholas J.
Xie, Jingyi
Yu, Rui
Lee, Sooyeon
Billah, Syed Masum
Carroll, John M.
Human-Computer Interaction
Assistive technologies for people with visual impairments (PVI) have made significant advancements, particularly with the integration of artificial intelligence (AI) and real-time sensor technologies. However, current solutions often require PVI to switch between multiple apps and tools for tasks like image recognition, navigation, and obstacle detection, which can hinder a seamless and efficient user experience. In this paper, we present NaviGPT, a high-fidelity prototype that integrates LiDAR-based obstacle detection, vibration feedback, and large language model (LLM) responses to provide a comprehensive and real-time navigation aid for PVI. Unlike existing applications such as Be My AI and Seeing AI, NaviGPT combines image recognition and contextual navigation guidance into a single system, offering continuous feedback on the user's surroundings without the need for app-switching. Meanwhile, NaviGPT compensates for the response delays of LLM by using location and sensor data, aiming to provide practical and efficient navigation support for PVI in dynamic environments.
title Enhancing the Travel Experience for People with Visual Impairments through Multimodal Interaction: NaviGPT, A Real-Time AI-Driven Mobile Navigation System
topic Human-Computer Interaction
url https://arxiv.org/abs/2410.04005