Navig-AI-tion: Navigation by Contextual AI and Spatial Audio
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913014874112000 |
|---|---|
| author | Lystbæk, Mathias N. Adams, Haley Ananda, Ranjith Kagathi Gonzalez, Eric J Ballan, Luca Wu, Qiuxuan Colaço, Andrea Tan, Peter Gonzalez-Franco, Mar |
| author_facet | Lystbæk, Mathias N. Adams, Haley Ananda, Ranjith Kagathi Gonzalez, Eric J Ballan, Luca Wu, Qiuxuan Colaço, Andrea Tan, Peter Gonzalez-Franco, Mar |
| contents | Audio-only walking navigation can leave users disoriented, relying on vague cardinal directions and lacking real-time environmental context, leading to frequent errors. To address this, we present a novel system that integrates a Vision Language Model (VLM) with a spatial audio cue. Our system extracts environmental landmarks to anchor navigation instructions and, crucially, provides a directional spatial audio signal when the user faces the wrong direction, indicating the precise turn direction. In a user study (n=12), the spatial audio cue with VLM reduced route deviations compared to both VLM-only and Google Maps (audio-only) baseline systems. Users reported that the spatial audio cue effectively supported orientation and that landmark-anchored instructions provided a better navigation experience over audio-only Google Maps. This work serves as an initial look at the utility of future audio-only navigation systems for incorporating directional cues, especially real-time corrective spatial audio. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_13200 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Navig-AI-tion: Navigation by Contextual AI and Spatial Audio Lystbæk, Mathias N. Adams, Haley Ananda, Ranjith Kagathi Gonzalez, Eric J Ballan, Luca Wu, Qiuxuan Colaço, Andrea Tan, Peter Gonzalez-Franco, Mar Human-Computer Interaction Audio-only walking navigation can leave users disoriented, relying on vague cardinal directions and lacking real-time environmental context, leading to frequent errors. To address this, we present a novel system that integrates a Vision Language Model (VLM) with a spatial audio cue. Our system extracts environmental landmarks to anchor navigation instructions and, crucially, provides a directional spatial audio signal when the user faces the wrong direction, indicating the precise turn direction. In a user study (n=12), the spatial audio cue with VLM reduced route deviations compared to both VLM-only and Google Maps (audio-only) baseline systems. Users reported that the spatial audio cue effectively supported orientation and that landmark-anchored instructions provided a better navigation experience over audio-only Google Maps. This work serves as an initial look at the utility of future audio-only navigation systems for incorporating directional cues, especially real-time corrective spatial audio. |
| title | Navig-AI-tion: Navigation by Contextual AI and Spatial Audio |
| topic | Human-Computer Interaction |
| url | https://arxiv.org/abs/2603.13200 |