MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yue, Xiang, Zheng, Tianyu, Ni, Yuansheng, Wang, Yubo, Zhang, Kai, Tong, Shengbang, Sun, Yuxuan, Yu, Botao, Zhang, Ge, Sun, Huan, Su, Yu, Chen, Wenhu, Neubig, Graham
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910960404398080
author Yue, Xiang
Zheng, Tianyu
Ni, Yuansheng
Wang, Yubo
Zhang, Kai
Tong, Shengbang
Sun, Yuxuan
Yu, Botao
Zhang, Ge
Sun, Huan
Su, Yu
Chen, Wenhu
Neubig, Graham
author_facet Yue, Xiang
Zheng, Tianyu
Ni, Yuansheng
Wang, Yubo
Zhang, Kai
Tong, Shengbang
Sun, Yuxuan
Yu, Botao
Zhang, Ge
Sun, Huan
Su, Yu
Chen, Wenhu
Neubig, Graham
contents This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly "see" and "read" simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Yue, Xiang
Zheng, Tianyu
Ni, Yuansheng
Wang, Yubo
Zhang, Kai
Tong, Shengbang
Sun, Yuxuan
Yu, Botao
Zhang, Ge
Sun, Huan
Su, Yu
Chen, Wenhu
Neubig, Graham
Computation and Language
Computer Vision and Pattern Recognition
This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly "see" and "read" simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.
title MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.02813