Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Meng, Iyer, Akhil, Pavel, Amy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909697962934272
author Chen, Meng
Iyer, Akhil
Pavel, Amy
author_facet Chen, Meng
Iyer, Akhil
Pavel, Amy
contents Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-checking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical. We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image. We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants. Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses. 14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15692
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
Chen, Meng
Iyer, Akhil
Pavel, Amy
Human-Computer Interaction
Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-checking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical. We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image. We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants. Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses. 14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media.
title Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
topic Human-Computer Interaction
url https://arxiv.org/abs/2507.15692