A Large Vision-Language Model based Environment Perception System for Visually Impaired People
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zezhou, Liu, Zhaoxiang, Wang, Kai, Wang, Kohou, Lian, Shiguo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
Vision-based Wearable Steering Assistance for People with Impaired Vision in Jogging
by: Liu, Xiaotong, et al.
Published: (2024)
by: Liu, Xiaotong, et al.
Published: (2024)
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)
by: Cui, Yajie, et al.
Published: (2023)
PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation
by: Fan, Qingyu, et al.
Published: (2026)
by: Fan, Qingyu, et al.
Published: (2026)
Piculet: Specialized Models-Guided Hallucination Decrease for MultiModal Large Language Models
by: Wang, Kohou, et al.
Published: (2024)
by: Wang, Kohou, et al.
Published: (2024)
Unified Vision-Language-Action Model
by: Wang, Yuqi, et al.
Published: (2025)
by: Wang, Yuqi, et al.
Published: (2025)
MemoNav: Working Memory Model for Visual Navigation
by: Li, Hongxin, et al.
Published: (2024)
by: Li, Hongxin, et al.
Published: (2024)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025)
by: Wang, Sheng
Published: (2025)
KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition
by: Liu, Zhaoxiang, et al.
Published: (2026)
by: Liu, Zhaoxiang, et al.
Published: (2026)
Towards Perception-based Collision Avoidance for UAVs when Guiding the Visually Impaired
by: Raj, Suman, et al.
Published: (2025)
by: Raj, Suman, et al.
Published: (2025)
MapDream: Task-Driven Map Learning for Vision-Language Navigation
by: Lian, Guoxin, et al.
Published: (2026)
by: Lian, Guoxin, et al.
Published: (2026)
PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
by: liu, Xiang, et al.
Published: (2025)
by: liu, Xiang, et al.
Published: (2025)
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
by: Ye, Jinhui, et al.
Published: (2026)
by: Ye, Jinhui, et al.
Published: (2026)
TP3M: Transformer-based Pseudo 3D Image Matching with Reference Image
by: Han, Liming, et al.
Published: (2024)
by: Han, Liming, et al.
Published: (2024)
FIReStereo: Forest InfraRed Stereo Dataset for UAS Depth Perception in Visually Degraded Environments
by: Dhrafani, Devansh, et al.
Published: (2024)
by: Dhrafani, Devansh, et al.
Published: (2024)
LIBERO-X: Robustness Litmus for Vision-Language-Action Models
by: Wang, Guodong, et al.
Published: (2026)
by: Wang, Guodong, et al.
Published: (2026)
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
by: Wang, Guankun, et al.
Published: (2024)
by: Wang, Guankun, et al.
Published: (2024)
SE-VLN: A Self-Evolving Vision-Language Navigation Framework Based on Multimodal Large Language Models
by: Dong, Xiangyu, et al.
Published: (2025)
by: Dong, Xiangyu, et al.
Published: (2025)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
by: Tian, Kefei, et al.
Published: (2026)
by: Tian, Kefei, et al.
Published: (2026)
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
by: Li, Anqi, et al.
Published: (2025)
by: Li, Anqi, et al.
Published: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
Language-Conditioned World Modeling for Visual Navigation
by: Dong, Yifei, et al.
Published: (2026)
by: Dong, Yifei, et al.
Published: (2026)
Rethinking the Embodied Gap in Vision-and-Language Navigation: A Holistic Study of Physical and Visual Disparities
by: Wang, Liuyi, et al.
Published: (2025)
by: Wang, Liuyi, et al.
Published: (2025)
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
by: He, Jiawei, et al.
Published: (2025)
by: He, Jiawei, et al.
Published: (2025)
An Embedded Real-time Object Alert System for Visually Impaired: A Monocular Depth Estimation based Approach through Computer Vision
by: Anjom, Jareen, et al.
Published: (2025)
by: Anjom, Jareen, et al.
Published: (2025)
ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
by: Wang, Xiyang, et al.
Published: (2026)
by: Wang, Xiyang, et al.
Published: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)
by: Yang, Sheng, et al.
Published: (2025)
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
by: Wei, Ziyu, et al.
Published: (2026)
by: Wei, Ziyu, et al.
Published: (2026)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
by: Zhang, Zezhou, et al.
Published: (2026)
by: Zhang, Zezhou, et al.
Published: (2026)
MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
by: Zhao, Kaikai, et al.
Published: (2025)
by: Zhao, Kaikai, et al.
Published: (2025)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video Generation
by: Gao, Huanlin, et al.
Published: (2025)
by: Gao, Huanlin, et al.
Published: (2025)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
by: Shen, Boyang, et al.
Published: (2026)
by: Shen, Boyang, et al.
Published: (2026)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
Similar Items
-
A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
by: Liu, Xiang, et al.
Published: (2025) -
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
by: Wang, Kohou, et al.
Published: (2025) -
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025) -
Vision-based Wearable Steering Assistance for People with Impaired Vision in Jogging
by: Liu, Xiaotong, et al.
Published: (2024) -
Patch-wise Auto-Encoder for Visual Anomaly Detection
by: Cui, Yajie, et al.
Published: (2023)