MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xiaolong, Kang, Zhaolu, Zhai, Wangyuxuan, Lou, Xinyue, Lai, Yunghwei, Wang, Ziyue, Wang, Yawen, Huang, Kaiyu, Wang, Yile, Li, Peng, Liu, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912607290523648
author Wang, Xiaolong
Kang, Zhaolu
Zhai, Wangyuxuan
Lou, Xinyue
Lai, Yunghwei
Wang, Ziyue
Wang, Yawen
Huang, Kaiyu
Wang, Yile
Li, Peng
Liu, Yang
author_facet Wang, Xiaolong
Kang, Zhaolu
Zhai, Wangyuxuan
Lou, Xinyue
Lai, Yunghwei
Wang, Ziyue
Wang, Yawen
Huang, Kaiyu
Wang, Yile
Li, Peng
Liu, Yang
contents Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text pairs with clear and explicit meanings. However, resolving the inherent ambiguities present in real-world language and visual contexts remains a challenge. Existing multimodal benchmarks typically overlook linguistic and visual ambiguities, relying mainly on unimodal context for disambiguation and thus failing to exploit the mutual clarification potential between modalities. To bridge this gap, we introduce MUCAR, a novel and challenging benchmark designed explicitly for evaluating multimodal ambiguity resolution across multilingual and cross-modal scenarios. MUCAR includes first a multilingual dataset where ambiguous textual expressions are uniquely resolved by corresponding visual contexts, and second a dual-ambiguity dataset that systematically pairs ambiguous images with ambiguous textual contexts, with each combination carefully constructed to yield a single, clear interpretation through mutual disambiguation. Extensive evaluations involving 19 state-of-the-art multimodal models--encompassing both open-source and proprietary architectures--reveal substantial gaps compared to human-level performance, highlighting the need for future research into more sophisticated cross-modal ambiguity comprehension methods, further pushing the boundaries of multimodal reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
Wang, Xiaolong
Kang, Zhaolu
Zhai, Wangyuxuan
Lou, Xinyue
Lai, Yunghwei
Wang, Ziyue
Wang, Yawen
Huang, Kaiyu
Wang, Yile
Li, Peng
Liu, Yang
Computation and Language
Machine Learning
Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text pairs with clear and explicit meanings. However, resolving the inherent ambiguities present in real-world language and visual contexts remains a challenge. Existing multimodal benchmarks typically overlook linguistic and visual ambiguities, relying mainly on unimodal context for disambiguation and thus failing to exploit the mutual clarification potential between modalities. To bridge this gap, we introduce MUCAR, a novel and challenging benchmark designed explicitly for evaluating multimodal ambiguity resolution across multilingual and cross-modal scenarios. MUCAR includes first a multilingual dataset where ambiguous textual expressions are uniquely resolved by corresponding visual contexts, and second a dual-ambiguity dataset that systematically pairs ambiguous images with ambiguous textual contexts, with each combination carefully constructed to yield a single, clear interpretation through mutual disambiguation. Extensive evaluations involving 19 state-of-the-art multimodal models--encompassing both open-source and proprietary architectures--reveal substantial gaps compared to human-level performance, highlighting the need for future research into more sophisticated cross-modal ambiguity comprehension methods, further pushing the boundaries of multimodal reasoning.
title MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.17046