mmWalk: Towards Multi-modal Multi-view Walking Assistance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ying, Kedi, Liu, Ruiping, Chen, Chongyan, Tao, Mingzhe, Shi, Hao, Yang, Kailun, Zhang, Jiaming, Stiefelhagen, Rainer
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911228374286336
author Ying, Kedi
Liu, Ruiping
Chen, Chongyan
Tao, Mingzhe
Shi, Hao
Yang, Kailun
Zhang, Jiaming
Stiefelhagen, Rainer
author_facet Ying, Kedi
Liu, Ruiping
Chen, Chongyan
Tao, Mingzhe
Shi, Hao
Yang, Kailun
Zhang, Jiaming
Stiefelhagen, Rainer
contents Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset that integrates multi-view sensor and accessibility-oriented features for outdoor safe navigation. Our dataset comprises 120 manually controlled, scenario-categorized walking trajectories with 62k synchronized frames. It contains over 559k panoramic images across RGB, depth, and semantic modalities. Furthermore, to emphasize real-world relevance, each trajectory involves outdoor corner cases and accessibility-specific landmarks for BLV users. Additionally, we generate mmWalkVQA, a VQA benchmark with over 69k visual question-answer triplets across 9 categories tailored for safe and informed walking assistance. We evaluate state-of-the-art Vision-Language Models (VLMs) using zero- and few-shot settings and found they struggle with our risk assessment and navigational tasks. We validate our mmWalk-finetuned model on real-world datasets and show the effectiveness of our dataset for advancing multi-modal walking assistance.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle mmWalk: Towards Multi-modal Multi-view Walking Assistance
Ying, Kedi
Liu, Ruiping
Chen, Chongyan
Tao, Mingzhe
Shi, Hao
Yang, Kailun
Zhang, Jiaming
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset that integrates multi-view sensor and accessibility-oriented features for outdoor safe navigation. Our dataset comprises 120 manually controlled, scenario-categorized walking trajectories with 62k synchronized frames. It contains over 559k panoramic images across RGB, depth, and semantic modalities. Furthermore, to emphasize real-world relevance, each trajectory involves outdoor corner cases and accessibility-specific landmarks for BLV users. Additionally, we generate mmWalkVQA, a VQA benchmark with over 69k visual question-answer triplets across 9 categories tailored for safe and informed walking assistance. We evaluate state-of-the-art Vision-Language Models (VLMs) using zero- and few-shot settings and found they struggle with our risk assessment and navigational tasks. We validate our mmWalk-finetuned model on real-world datasets and show the effectiveness of our dataset for advancing multi-modal walking assistance.
title mmWalk: Towards Multi-modal Multi-view Walking Assistance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11520