MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zhaowei, Yu, Wenhao, Ren, Xiyu, Zhang, Jipeng, Zhao, Yu, Saxena, Rohit, Cheng, Liang, Wong, Ginny, See, Simon, Minervini, Pasquale, Song, Yangqiu, Steedman, Mark
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911193631817728
author Wang, Zhaowei
Yu, Wenhao
Ren, Xiyu
Zhang, Jipeng
Zhao, Yu
Saxena, Rohit
Cheng, Liang
Wong, Ginny
See, Simon
Minervini, Pasquale
Song, Yangqiu
Steedman, Mark
author_facet Wang, Zhaowei
Yu, Wenhao
Ren, Xiyu
Zhang, Jipeng
Zhao, Yu
Saxena, Rohit
Cheng, Liang
Wong, Ginny
See, Simon
Minervini, Pasquale
Song, Yangqiu
Steedman, Mark
contents The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
Wang, Zhaowei
Yu, Wenhao
Ren, Xiyu
Zhang, Jipeng
Zhao, Yu
Saxena, Rohit
Cheng, Liang
Wong, Ginny
See, Simon
Minervini, Pasquale
Song, Yangqiu
Steedman, Mark
Computer Vision and Pattern Recognition
Computation and Language
The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a single forward pass. In this work, we introduce MMLongBench, the first benchmark covering a diverse set of long-context vision-language tasks, to evaluate LCVLMs effectively and thoroughly. MMLongBench is composed of 13,331 examples spanning five different categories of downstream tasks, such as Visual RAG and Many-Shot ICL. It also provides broad coverage of image types, including various natural and synthetic images. To assess the robustness of the models to different input lengths, all examples are delivered at five standardized input lengths (8K-128K tokens) via a cross-modal tokenization scheme that combines vision patches and text tokens. Through a thorough benchmarking of 46 closed-source and open-source LCVLMs, we provide a comprehensive analysis of the current models' vision-language long-context ability. Our results show that: i) performance on a single task is a weak proxy for overall long-context capability; ii) both closed-source and open-source models face challenges in long-context vision-language tasks, indicating substantial room for future improvement; iii) models with stronger reasoning ability tend to exhibit better long-context performance. By offering wide task coverage, various image types, and rigorous length control, MMLongBench provides the missing foundation for diagnosing and advancing the next generation of LCVLMs.
title MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.10610