R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chunyi, Zhang, Jianbo, Zhang, Zicheng, Wu, Haoning, Tian, Yuan, Sun, Wei, Lu, Guo, Liu, Xiaohong, Min, Xiongkuo, Lin, Weisi, Zhai, Guangtao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912062792269824
author Li, Chunyi
Zhang, Jianbo
Zhang, Zicheng
Wu, Haoning
Tian, Yuan
Sun, Wei
Lu, Guo
Liu, Xiaohong
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
author_facet Li, Chunyi
Zhang, Jianbo
Zhang, Zicheng
Wu, Haoning
Tian, Yuan
Sun, Wei
Lu, Guo
Liu, Xiaohong
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
contents The outstanding performance of Large Multimodal Models (LMMs) has made them widely applied in vision-related tasks. However, various corruptions in the real world mean that images will not be as ideal as in simulations, presenting significant challenges for the practical application of LMMs. To address this issue, we introduce R-Bench, a benchmark focused on the **Real-world Robustness of LMMs**. Specifically, we: (a) model the complete link from user capture to LMMs reception, comprising 33 corruption dimensions, including 7 steps according to the corruption sequence, and 7 groups based on low-level attributes; (b) collect reference/distorted image dataset before/after corruption, including 2,970 question-answer pairs with human labeling; (c) propose comprehensive evaluation for absolute/relative robustness and benchmark 20 mainstream LMMs. Results show that while LMMs can correctly handle the original reference images, their performance is not stable when faced with distorted images, and there is a significant gap in robustness compared to the human visual system. We hope that R-Bench will inspire improving the robustness of LMMs, **extending them from experimental simulations to the real-world application**. Check https://q-future.github.io/R-Bench for details.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05474
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
Li, Chunyi
Zhang, Jianbo
Zhang, Zicheng
Wu, Haoning
Tian, Yuan
Sun, Wei
Lu, Guo
Liu, Xiaohong
Min, Xiongkuo
Lin, Weisi
Zhai, Guangtao
Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
The outstanding performance of Large Multimodal Models (LMMs) has made them widely applied in vision-related tasks. However, various corruptions in the real world mean that images will not be as ideal as in simulations, presenting significant challenges for the practical application of LMMs. To address this issue, we introduce R-Bench, a benchmark focused on the **Real-world Robustness of LMMs**. Specifically, we: (a) model the complete link from user capture to LMMs reception, comprising 33 corruption dimensions, including 7 steps according to the corruption sequence, and 7 groups based on low-level attributes; (b) collect reference/distorted image dataset before/after corruption, including 2,970 question-answer pairs with human labeling; (c) propose comprehensive evaluation for absolute/relative robustness and benchmark 20 mainstream LMMs. Results show that while LMMs can correctly handle the original reference images, their performance is not stable when faced with distorted images, and there is a significant gap in robustness compared to the human visual system. We hope that R-Bench will inspire improving the robustness of LMMs, **extending them from experimental simulations to the real-world application**. Check https://q-future.github.io/R-Bench for details.
title R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
topic Computer Vision and Pattern Recognition
Multimedia
Image and Video Processing
url https://arxiv.org/abs/2410.05474