Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chenxu, Wang, Zhicai, Sheng, Yuan, Zhu, Xingyu, Hao, Yanbin, Wang, Xiang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917079419977728
author Li, Chenxu
Wang, Zhicai
Sheng, Yuan
Zhu, Xingyu
Hao, Yanbin
Wang, Xiang
author_facet Li, Chenxu
Wang, Zhicai
Sheng, Yuan
Zhu, Xingyu
Hao, Yanbin
Wang, Xiang
contents Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce \textbf{Res-Bench}, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16926
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
Li, Chenxu
Wang, Zhicai
Sheng, Yuan
Zhu, Xingyu
Hao, Yanbin
Wang, Xiang
Computer Vision and Pattern Recognition
Computation and Language
Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce \textbf{Res-Bench}, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement.
title Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2510.16926