RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Fei, Lu, Chengqiang, Shen, Yufan, Wang, Qimeng, Qian, Yicheng, Zhang, Haoxin, Gao, Yan, Wu, Yi, Hu, Yao, Wu, Zhen, Xing, Shangyu, Dai, Xinyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909799488159744
author Zhao, Fei
Lu, Chengqiang
Shen, Yufan
Wang, Qimeng
Qian, Yicheng
Zhang, Haoxin
Gao, Yan
Wu, Yi
Hu, Yao
Wu, Zhen
Xing, Shangyu
Dai, Xinyu
author_facet Zhao, Fei
Lu, Chengqiang
Shen, Yufan
Wang, Qimeng
Qian, Yicheng
Zhang, Haoxin
Gao, Yan
Wu, Yi
Hu, Yao
Wu, Zhen
Xing, Shangyu
Dai, Xinyu
contents While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first Chinese multimodal multi-image dataset, which contains 9393 samples and 69910 images. RealBench distinguishes itself by incorporating real user-generated content, ensuring high relevance to real-world applications. Additionally, the dataset covers a wide variety of scenes, image resolutions, and image structures, further increasing the difficulty of multi-image understanding. Ultimately, we conduct a comprehensive evaluation of RealBench using 21 multimodal LLMs of different sizes, including closed-source models that support multi-image inputs as well as open-source visual and video models. The experimental results indicate that even the most powerful closed-source models still face challenges when handling multi-image Chinese scenarios. Moreover, there remains a noticeable performance gap of around 71.8\% on average between open-source visual/video models and closed-source models. These results show that RealBench provides an important research foundation for further exploring multi-image understanding capabilities in the Chinese context.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
Zhao, Fei
Lu, Chengqiang
Shen, Yufan
Wang, Qimeng
Qian, Yicheng
Zhang, Haoxin
Gao, Yan
Wu, Yi
Hu, Yao
Wu, Zhen
Xing, Shangyu
Dai, Xinyu
Computation and Language
Multimedia
While various multimodal multi-image evaluation datasets have been emerged, but these datasets are primarily based on English, and there has yet to be a Chinese multi-image dataset. To fill this gap, we introduce RealBench, the first Chinese multimodal multi-image dataset, which contains 9393 samples and 69910 images. RealBench distinguishes itself by incorporating real user-generated content, ensuring high relevance to real-world applications. Additionally, the dataset covers a wide variety of scenes, image resolutions, and image structures, further increasing the difficulty of multi-image understanding. Ultimately, we conduct a comprehensive evaluation of RealBench using 21 multimodal LLMs of different sizes, including closed-source models that support multi-image inputs as well as open-source visual and video models. The experimental results indicate that even the most powerful closed-source models still face challenges when handling multi-image Chinese scenarios. Moreover, there remains a noticeable performance gap of around 71.8\% on average between open-source visual/video models and closed-source models. These results show that RealBench provides an important research foundation for further exploring multi-image understanding capabilities in the Chinese context.
title RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
topic Computation and Language
Multimedia
url https://arxiv.org/abs/2509.17421