PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Junyuan, Song, Jiahe, Wu, Jiang, Zhu, Runchuan, Shen, Guanlin, Wang, Shasha, Wei, Xingjian, Yang, Haote, Zhang, Songyang, Li, Weijia, Wang, Bin, Lin, Dahua, Wu, Lijun, He, Conghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915713412759552
author Gao, Junyuan
Song, Jiahe
Wu, Jiang
Zhu, Runchuan
Shen, Guanlin
Wang, Shasha
Wei, Xingjian
Yang, Haote
Zhang, Songyang
Li, Weijia
Wang, Bin
Lin, Dahua
Wu, Lijun
He, Conghui
author_facet Gao, Junyuan
Song, Jiahe
Wu, Jiang
Zhu, Runchuan
Shen, Guanlin
Wang, Shasha
Wei, Xingjian
Yang, Haote
Zhang, Songyang
Li, Weijia
Wang, Bin
Lin, Dahua
Wu, Lijun
He, Conghui
contents While Large Vision-Language Models (LVLMs) demonstrate promising multilingual capabilities, their evaluation is currently hindered by two critical limitations: (1) the use of non-parallel corpora, which conflates inherent language capability gaps with dataset artifacts, precluding a fair assessment of cross-lingual alignment; and (2) disjointed multimodal inputs, which deviate from real-world scenarios where most texts are embedded within visual contexts. To address these challenges, we propose PM4Bench, the first Multilingual Multi-Modal Multi-task Benchmark constructed on a strictly parallel corpus across 10 languages. By eliminating content divergence, our benchmark enables a fair comparison of model capabilities across different languages. We also introduce a vision setting where textual queries are visually fused into images, compelling models to jointly "see," "read," and "think". Extensive evaluation of 10 LVLMs uncover a substantial performance drop in the Vision setting compared to standard inputs. Further analysis reveals that OCR capability is not only a general bottleneck but also contributes to cross-lingual performance disparities, suggesting that improving multilingual OCR is essential for advancing LVLM performance. We will release PM4Bench at https://github.com/opendatalab/PM4Bench .
format Preprint
id arxiv_https___arxiv_org_abs_2503_18484
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
Gao, Junyuan
Song, Jiahe
Wu, Jiang
Zhu, Runchuan
Shen, Guanlin
Wang, Shasha
Wei, Xingjian
Yang, Haote
Zhang, Songyang
Li, Weijia
Wang, Bin
Lin, Dahua
Wu, Lijun
He, Conghui
Computer Vision and Pattern Recognition
Computation and Language
While Large Vision-Language Models (LVLMs) demonstrate promising multilingual capabilities, their evaluation is currently hindered by two critical limitations: (1) the use of non-parallel corpora, which conflates inherent language capability gaps with dataset artifacts, precluding a fair assessment of cross-lingual alignment; and (2) disjointed multimodal inputs, which deviate from real-world scenarios where most texts are embedded within visual contexts. To address these challenges, we propose PM4Bench, the first Multilingual Multi-Modal Multi-task Benchmark constructed on a strictly parallel corpus across 10 languages. By eliminating content divergence, our benchmark enables a fair comparison of model capabilities across different languages. We also introduce a vision setting where textual queries are visually fused into images, compelling models to jointly "see," "read," and "think". Extensive evaluation of 10 LVLMs uncover a substantial performance drop in the Vision setting compared to standard inputs. Further analysis reveals that OCR capability is not only a general bottleneck but also contributes to cross-lingual performance disparities, suggesting that improving multilingual OCR is essential for advancing LVLM performance. We will release PM4Bench at https://github.com/opendatalab/PM4Bench .
title PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.18484