Navig-AI-tion: Navigation by Contextual AI and Spatial Audio

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lystbæk, Mathias N., Adams, Haley, Ananda, Ranjith Kagathi, Gonzalez, Eric J, Ballan, Luca, Wu, Qiuxuan, Colaço, Andrea, Tan, Peter, Gonzalez-Franco, Mar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913014874112000
author Lystbæk, Mathias N.
Adams, Haley
Ananda, Ranjith Kagathi
Gonzalez, Eric J
Ballan, Luca
Wu, Qiuxuan
Colaço, Andrea
Tan, Peter
Gonzalez-Franco, Mar
author_facet Lystbæk, Mathias N.
Adams, Haley
Ananda, Ranjith Kagathi
Gonzalez, Eric J
Ballan, Luca
Wu, Qiuxuan
Colaço, Andrea
Tan, Peter
Gonzalez-Franco, Mar
contents Audio-only walking navigation can leave users disoriented, relying on vague cardinal directions and lacking real-time environmental context, leading to frequent errors. To address this, we present a novel system that integrates a Vision Language Model (VLM) with a spatial audio cue. Our system extracts environmental landmarks to anchor navigation instructions and, crucially, provides a directional spatial audio signal when the user faces the wrong direction, indicating the precise turn direction. In a user study (n=12), the spatial audio cue with VLM reduced route deviations compared to both VLM-only and Google Maps (audio-only) baseline systems. Users reported that the spatial audio cue effectively supported orientation and that landmark-anchored instructions provided a better navigation experience over audio-only Google Maps. This work serves as an initial look at the utility of future audio-only navigation systems for incorporating directional cues, especially real-time corrective spatial audio.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13200
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Navig-AI-tion: Navigation by Contextual AI and Spatial Audio
Lystbæk, Mathias N.
Adams, Haley
Ananda, Ranjith Kagathi
Gonzalez, Eric J
Ballan, Luca
Wu, Qiuxuan
Colaço, Andrea
Tan, Peter
Gonzalez-Franco, Mar
Human-Computer Interaction
Audio-only walking navigation can leave users disoriented, relying on vague cardinal directions and lacking real-time environmental context, leading to frequent errors. To address this, we present a novel system that integrates a Vision Language Model (VLM) with a spatial audio cue. Our system extracts environmental landmarks to anchor navigation instructions and, crucially, provides a directional spatial audio signal when the user faces the wrong direction, indicating the precise turn direction. In a user study (n=12), the spatial audio cue with VLM reduced route deviations compared to both VLM-only and Google Maps (audio-only) baseline systems. Users reported that the spatial audio cue effectively supported orientation and that landmark-anchored instructions provided a better navigation experience over audio-only Google Maps. This work serves as an initial look at the utility of future audio-only navigation systems for incorporating directional cues, especially real-time corrective spatial audio.
title Navig-AI-tion: Navigation by Contextual AI and Spatial Audio
topic Human-Computer Interaction
url https://arxiv.org/abs/2603.13200