WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Zhiqiang, Zhang, Ting, Deng, Ying, Zhang, Jiapei, Zhu, Yeshuang, Jia, Zexi, Zhou, Jie, Zhang, Jinchao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912257482424320
author Yuan, Zhiqiang
Zhang, Ting
Deng, Ying
Zhang, Jiapei
Zhu, Yeshuang
Jia, Zexi
Zhou, Jie
Zhang, Jinchao
author_facet Yuan, Zhiqiang
Zhang, Ting
Deng, Ying
Zhang, Jiapei
Zhu, Yeshuang
Jia, Zexi
Zhou, Jie
Zhang, Jinchao
contents Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has become popular. However, the existing methods of walking guidance are mainly based on self-curated question-answering datasets that are not publicly accessible, without a standardized benchmark for training or evaluation. Moreover, walking assistance often requires real-time streaming video analysis and the generation of concise yet informative reminders, making VLMs struggle due to excessive responses and low efficiency in inferences. In this paper, we introduce the first large-scale dataset dedicated to walking assistance, comprising 12,000 video-annotation pairs, to provide a unified benchmark for training and evaluating systems to help visually-impaired individuals walk. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code are available at https://walkvlm2024.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20903
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
Yuan, Zhiqiang
Zhang, Ting
Deng, Ying
Zhang, Jiapei
Zhu, Yeshuang
Jia, Zexi
Zhou, Jie
Zhang, Jinchao
Computer Vision and Pattern Recognition
Artificial Intelligence
Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has become popular. However, the existing methods of walking guidance are mainly based on self-curated question-answering datasets that are not publicly accessible, without a standardized benchmark for training or evaluation. Moreover, walking assistance often requires real-time streaming video analysis and the generation of concise yet informative reminders, making VLMs struggle due to excessive responses and low efficiency in inferences. In this paper, we introduce the first large-scale dataset dedicated to walking assistance, comprising 12,000 video-annotation pairs, to provide a unified benchmark for training and evaluating systems to help visually-impaired individuals walk. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporal-aware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code are available at https://walkvlm2024.github.io.
title WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.20903