@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909402533986304 |
|---|---|
| author | Jiang, Xin Zheng, Junwei Liu, Ruiping Li, Jiahang Zhang, Jiaming Matthiesen, Sven Stiefelhagen, Rainer |
| author_facet | Jiang, Xin Zheng, Junwei Liu, Ruiping Li, Jiahang Zhang, Jiaming Matthiesen, Sven Stiefelhagen, Rainer |
| contents | As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_14215 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | @Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology Jiang, Xin Zheng, Junwei Liu, Ruiping Li, Jiahang Zhang, Jiaming Matthiesen, Sven Stiefelhagen, Rainer Computer Vision and Pattern Recognition As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework. |
| title | @Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.14215 |