Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kang, Wan Ju, Kim, Eunki, An, Na Min, Kim, Sangryul, Choi, Haemin, Kwak, Ki Hoon, Thorne, James
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913741079052288
author Kang, Wan Ju
Kim, Eunki
An, Na Min
Kim, Sangryul
Choi, Haemin
Kwak, Ki Hoon
Thorne, James
author_facet Kang, Wan Ju
Kim, Eunki
An, Na Min
Kim, Sangryul
Choi, Haemin
Kwak, Ki Hoon
Thorne, James
contents Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe visuals with ease, but existing studies have shown that direct generations by them are costly, bias-prone, and somewhat lacking by BLV standards. In this study, we ask sighted individuals to assess -- rather than produce -- diagram descriptions generated by vision-language models (VLM) that have been guided with latent supervision via a multi-pass inference. The sighted assessments prove effective and useful to professional educators who are themselves BLV and teach visually impaired learners. We release Sightation, a collection of diagram description datasets spanning 5k diagrams and 137k samples for completion, preference, retrieval, question answering, and reasoning training purposes and demonstrate their fine-tuning potential in various downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
Kang, Wan Ju
Kim, Eunki
An, Na Min
Kim, Sangryul
Choi, Haemin
Kwak, Ki Hoon
Thorne, James
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe visuals with ease, but existing studies have shown that direct generations by them are costly, bias-prone, and somewhat lacking by BLV standards. In this study, we ask sighted individuals to assess -- rather than produce -- diagram descriptions generated by vision-language models (VLM) that have been guided with latent supervision via a multi-pass inference. The sighted assessments prove effective and useful to professional educators who are themselves BLV and teach visually impaired learners. We release Sightation, a collection of diagram description datasets spanning 5k diagrams and 137k samples for completion, preference, retrieval, question answering, and reasoning training purposes and demonstrate their fine-tuning potential in various downstream tasks.
title Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
topic Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2503.13369