@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Xin, Zheng, Junwei, Liu, Ruiping, Li, Jiahang, Zhang, Jiaming, Matthiesen, Sven, Stiefelhagen, Rainer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909402533986304
author Jiang, Xin
Zheng, Junwei
Liu, Ruiping
Li, Jiahang
Zhang, Jiaming
Matthiesen, Sven
Stiefelhagen, Rainer
author_facet Jiang, Xin
Zheng, Junwei
Liu, Ruiping
Li, Jiahang
Zhang, Jiaming
Matthiesen, Sven
Stiefelhagen, Rainer
contents As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14215
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle @Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
Jiang, Xin
Zheng, Junwei
Liu, Ruiping
Li, Jiahang
Zhang, Jiaming
Matthiesen, Sven
Stiefelhagen, Rainer
Computer Vision and Pattern Recognition
As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework.
title @Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.14215