Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
Fuente:
arXiv
Saved in:
| Main Authors: | Natalie, Rosiana, Xu, Wenqian, Chang, Ruei-Che, Mihalcea, Rada, Guo, Anhong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
by: Chang, Ruei-Che, et al.
Published: (2025)
by: Chang, Ruei-Che, et al.
Published: (2025)
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2026)
by: Chang, Ruei-Che, et al.
Published: (2026)
Audio Description Customization
by: Natalie, Rosiana, et al.
Published: (2024)
by: Natalie, Rosiana, et al.
Published: (2024)
WorldScribe: Towards Context-Aware Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
by: Chang, Ruei-Che, et al.
Published: (2026)
by: Chang, Ruei-Che, et al.
Published: (2026)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
ObjectFinder: An Open-Vocabulary Assistive System for Interactive Object Search by Blind People
by: Liu, Ruiping, et al.
Published: (2024)
by: Liu, Ruiping, et al.
Published: (2024)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
by: Verma, Arnav, et al.
Published: (2025)
by: Verma, Arnav, et al.
Published: (2025)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
by: Nowicki, Filip, et al.
Published: (2026)
by: Nowicki, Filip, et al.
Published: (2026)
An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
by: Huang, Yifei, et al.
Published: (2025)
by: Huang, Yifei, et al.
Published: (2025)
Vision Language Models as Values Detectors
by: Abbo, Giulio Antonio, et al.
Published: (2025)
by: Abbo, Giulio Antonio, et al.
Published: (2025)
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection
by: Chavan, Kunal, et al.
Published: (2025)
by: Chavan, Kunal, et al.
Published: (2025)
iTrace: Click-Based Gaze Visualization on the Apple Vision Pro
by: Mehmedova, Esra, et al.
Published: (2025)
by: Mehmedova, Esra, et al.
Published: (2025)
Real-Time Cellist Postural Evaluation With On-Device Computer Vision
by: Wang, Paolo, et al.
Published: (2026)
by: Wang, Paolo, et al.
Published: (2026)
"It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
by: Garg, Kapil, et al.
Published: (2025)
by: Garg, Kapil, et al.
Published: (2025)
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
by: Islam, Md Touhidul, et al.
Published: (2024)
by: Islam, Md Touhidul, et al.
Published: (2024)
VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Improve accessibility for Low Vision and Blind people using Machine Learning and Computer Vision
by: Shukurov, Jasur
Published: (2024)
by: Shukurov, Jasur
Published: (2024)
Do Vision Language Models Understand Human Engagement in Games?
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
by: Penuela, Ricardo Gonzalez, et al.
Published: (2025)
HarassGuard: Detecting Harassment Behaviors in Social Virtual Reality with Vision-Language Models
by: Lee, Junhee, et al.
Published: (2026)
by: Lee, Junhee, et al.
Published: (2026)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
ScreenAgent: A Vision Language Model-driven Computer Control Agent
by: Niu, Runliang, et al.
Published: (2024)
by: Niu, Runliang, et al.
Published: (2024)
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble
by: Duan, Lin, et al.
Published: (2025)
by: Duan, Lin, et al.
Published: (2025)
VisionCAD: An Integration-Free Radiology Copilot Framework
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
by: Hernandez, Juan Manuel, et al.
Published: (2026)
by: Hernandez, Juan Manuel, et al.
Published: (2026)
VFA: Vision Frequency Analysis of Foundation Models and Human
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2024)
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2024)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
by: Mazumdar, Amrita, et al.
Published: (2026)
by: Mazumdar, Amrita, et al.
Published: (2026)
Computer Vision for Objects used in Group Work: Challenges and Opportunities
by: Jung, Changsoo, et al.
Published: (2025)
by: Jung, Changsoo, et al.
Published: (2025)
Machine Vision-Based Surgical Lighting System:Design and Implementation
by: Gharghabi, Amir, et al.
Published: (2025)
by: Gharghabi, Amir, et al.
Published: (2025)
Weak-Annotation of HAR Datasets using Vision Foundation Models
by: Bock, Marius, et al.
Published: (2024)
by: Bock, Marius, et al.
Published: (2024)
Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models
by: Gallardo, Rodrigo, et al.
Published: (2025)
by: Gallardo, Rodrigo, et al.
Published: (2025)
SymbolSight: Minimizing Inter-Symbol Interference for Reading with Prosthetic Vision
by: Lesner, Jasmine, et al.
Published: (2026)
by: Lesner, Jasmine, et al.
Published: (2026)
MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
by: Mukhopadhyay, Srija, et al.
Published: (2024)
by: Mukhopadhyay, Srija, et al.
Published: (2024)
Essential, Yet Overlooked: Identity Verification Barriers for Blind and Low Vision People in Government Services
by: Oommen, Ryan John, et al.
Published: (2026)
by: Oommen, Ryan John, et al.
Published: (2026)
ClickAIXR: On-Device Multimodal Vision-Language Interaction with Real-World Objects in Extended Reality
by: Khan, Dawar, et al.
Published: (2026)
by: Khan, Dawar, et al.
Published: (2026)
Similar Items
-
Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
by: Chang, Ruei-Che, et al.
Published: (2025) -
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2026) -
Audio Description Customization
by: Natalie, Rosiana, et al.
Published: (2024) -
WorldScribe: Towards Context-Aware Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2024) -
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024)