A Multi-Modal Foundation Model to Assist People with Blindness and Low Vision in Environmental Interaction
Fuente:
arXiv
Guardado en:
| Autores principales: | Hao, Yu, Yang, Fan, Huang, Hao, Yuan, Shuaihang, Rangan, Sundeep, Rizzo, John-Ross, Wang, Yao, Fang, Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
por: Magay, Alexey, et al.
Publicado: (2025)
por: Magay, Alexey, et al.
Publicado: (2025)
A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models
por: Huang, Hao, et al.
Publicado: (2025)
por: Huang, Hao, et al.
Publicado: (2025)
Meta-Learning 3D Shape Segmentation Functions
por: Hao, Yu, et al.
Publicado: (2021)
por: Hao, Yu, et al.
Publicado: (2021)
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
por: Yuan, Shuaihang, et al.
Publicado: (2024)
por: Yuan, Shuaihang, et al.
Publicado: (2024)
Learning-Based Signal Recovery in Nonlinear Systems with Spectrally Separated Interference
por: Joy, Jayadev, et al.
Publicado: (2026)
por: Joy, Jayadev, et al.
Publicado: (2026)
Computationally Efficient Signal Detection with Unknown Bandwidths
por: Rasteh, Ali, et al.
Publicado: (2025)
por: Rasteh, Ali, et al.
Publicado: (2025)
A Multimodal Assistive System for Product Localization and Retrieval for People who are Blind or have Low Vision
por: Ruan, Ligao, et al.
Publicado: (2026)
por: Ruan, Ligao, et al.
Publicado: (2026)
Beyond Point Estimates: Likelihood-Based Full-Posterior Wireless Localization
por: Lei, Haozhe, et al.
Publicado: (2025)
por: Lei, Haozhe, et al.
Publicado: (2025)
Zero-shot Object Navigation with Vision-Language Models Reasoning
por: Wen, Congcong, et al.
Publicado: (2024)
por: Wen, Congcong, et al.
Publicado: (2024)
Exploring the Reliability of Foundation Model-Based Frontier Selection in Zero-Shot Object Goal Navigation
por: Yuan, Shuaihang, et al.
Publicado: (2024)
por: Yuan, Shuaihang, et al.
Publicado: (2024)
Low-rank Preconditioning in Beamspace Domain For Massive MU-MIMO Long-Term Beamforming
por: Kiani, Amirreza, et al.
Publicado: (2026)
por: Kiani, Amirreza, et al.
Publicado: (2026)
Residual Gaze Behavior During Navigation in Blindness and Low Vision
por: Feng, Junchi, et al.
Publicado: (2025)
por: Feng, Junchi, et al.
Publicado: (2025)
Long-Form Answers to Visual Questions from Blind and Low Vision People
por: Huh, Mina, et al.
Publicado: (2024)
por: Huh, Mina, et al.
Publicado: (2024)
Scalable Long-Term Beamforming for Massive Multi-User MIMO
por: Rasteh, Ali, et al.
Publicado: (2025)
por: Rasteh, Ali, et al.
Publicado: (2025)
Channel Modeling for FR3 Upper Mid-band via Generative Adversarial Networks
por: Hu, Yaqi, et al.
Publicado: (2024)
por: Hu, Yaqi, et al.
Publicado: (2024)
Integrating Retrospective Framework in Multi-Robot Collaboration
por: Liang, Jiazhao, et al.
Publicado: (2025)
por: Liang, Jiazhao, et al.
Publicado: (2025)
How Can Haptic Feedback Assist People with Blind and Low Vision (BLV): A Systematic Literature Review
por: Jiang, Chutian, et al.
Publicado: (2024)
por: Jiang, Chutian, et al.
Publicado: (2024)
AnyImageNav: Any-View Geometry for Precise Last-Meter Image-Goal Navigation
por: Deng, Yijie, et al.
Publicado: (2026)
por: Deng, Yijie, et al.
Publicado: (2026)
Near-Field Measurement System for the Upper Mid-Band
por: Rasteh, Ali, et al.
Publicado: (2024)
por: Rasteh, Ali, et al.
Publicado: (2024)
Transformer-Based Rate Prediction for Multi-Band Cellular Handsets
por: Chen, Ruibin, et al.
Publicado: (2025)
por: Chen, Ruibin, et al.
Publicado: (2025)
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
por: Wen, Congcong, et al.
Publicado: (2025)
por: Wen, Congcong, et al.
Publicado: (2025)
Reliable Semantic Understanding for Real World Zero-shot Object Goal Navigation
por: Unlu, Halil Utku, et al.
Publicado: (2024)
por: Unlu, Halil Utku, et al.
Publicado: (2024)
Haptics-based, higher-order Sensory Substitution designed for Object Negotiation in Blindness and Low Vision: Virtual Whiskers
por: Feng, Junchi, et al.
Publicado: (2024)
por: Feng, Junchi, et al.
Publicado: (2024)
Does Embodiment Matter to Biomechanics and Function? A Comparative Analysis of Head-Mounted and Hand-Held Assistive Devices for Individuals with Blindness and Low Vision
por: Seth, Gaurav, et al.
Publicado: (2025)
por: Seth, Gaurav, et al.
Publicado: (2025)
Upper Mid-Band Spectrum for 6G: Vision, Opportunity and Challenges
por: Bazzi, Ahmad, et al.
Publicado: (2025)
por: Bazzi, Ahmad, et al.
Publicado: (2025)
Advancing Network Digital Twin Framework for Generating Realistic Datasets
por: Stenhammar, Oscar, et al.
Publicado: (2026)
por: Stenhammar, Oscar, et al.
Publicado: (2026)
Terrestrial-Satellite Spectrum Sharing in the Upper Mid-Band with Interference Nulling
por: Kang, Seongjoon, et al.
Publicado: (2023)
por: Kang, Seongjoon, et al.
Publicado: (2023)
Interference Suppression for Massive MU-MIMO Long-Term Beamforming with Matrix Inversion Approximation
por: Kiani, Amirreza, et al.
Publicado: (2026)
por: Kiani, Amirreza, et al.
Publicado: (2026)
An Experimental Multi-Band Channel Characterization in the Upper Mid-Band
por: Bomfin, Roberto, et al.
Publicado: (2024)
por: Bomfin, Roberto, et al.
Publicado: (2024)
Wavelet Policy: Lifting Scheme for Policy Learning in Long-Horizon Tasks
por: Huang, Hao, et al.
Publicado: (2025)
por: Huang, Hao, et al.
Publicado: (2025)
Accessibility and Social Inclusivity: A Literature Review of Music Technology for Blind and Low Vision People
por: Zhang, Shumeng, et al.
Publicado: (2025)
por: Zhang, Shumeng, et al.
Publicado: (2025)
A Preliminary Assessment of Midhaul Links at 140 GHz using Ray-Tracing
por: Chintareddy, Sravan Reddy, et al.
Publicado: (2026)
por: Chintareddy, Sravan Reddy, et al.
Publicado: (2026)
Estimation of embedding vectors in high dimensions
por: Azar, Golara Ahmadi, et al.
Publicado: (2023)
por: Azar, Golara Ahmadi, et al.
Publicado: (2023)
Essential, Yet Overlooked: Identity Verification Barriers for Blind and Low Vision People in Government Services
por: Oommen, Ryan John, et al.
Publicado: (2026)
por: Oommen, Ryan John, et al.
Publicado: (2026)
Digital Twin-Enhanced Wireless Indoor Navigation: Achieving Efficient Environment Sensing with Zero-Shot Reinforcement Learning
por: Li, Tao, et al.
Publicado: (2023)
por: Li, Tao, et al.
Publicado: (2023)
FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech
por: Lai, Yuzhi, et al.
Publicado: (2025)
por: Lai, Yuzhi, et al.
Publicado: (2025)
Multi-faceted Sensory Substitution for Curb Alerting: A Pilot Investigation in Persons with Blindness and Low Vision
por: Ruan, Ligao, et al.
Publicado: (2024)
por: Ruan, Ligao, et al.
Publicado: (2024)
Generating Contextually-Relevant Navigation Instructions for Blind and Low Vision People
por: Merchant, Zain, et al.
Publicado: (2024)
por: Merchant, Zain, et al.
Publicado: (2024)
RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
por: Cao, Xinyun, et al.
Publicado: (2025)
por: Cao, Xinyun, et al.
Publicado: (2025)
Ejemplares similares
-
Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
por: Li, Yu, et al.
Publicado: (2026) -
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
por: Magay, Alexey, et al.
Publicado: (2025) -
A Chain-of-Thought Subspace Meta-Learning for Few-shot Image Captioning with Large Vision and Language Models
por: Huang, Hao, et al.
Publicado: (2025) -
Meta-Learning 3D Shape Segmentation Functions
por: Hao, Yu, et al.
Publicado: (2021) -
GAMap: Zero-Shot Object Goal Navigation with Multi-Scale Geometric-Affordance Guidance
por: Yuan, Shuaihang, et al.
Publicado: (2024)