MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fu, Chaoyou, Chen, Peixian, Shen, Yunhang, Qin, Yulei, Zhang, Mengdan, Lin, Xu, Yang, Jinrui, Zheng, Xiawu, Li, Ke, Sun, Xing, Wu, Yunsheng, Ji, Rongrong, Shan, Caifeng, He, Ran
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911228494872576
author Fu, Chaoyou
Chen, Peixian
Shen, Yunhang
Qin, Yulei
Zhang, Mengdan
Lin, Xu
Yang, Jinrui
Zheng, Xiawu
Li, Ke
Sun, Xing
Wu, Yunsheng
Ji, Rongrong
Shan, Caifeng
He, Ran
author_facet Fu, Chaoyou
Chen, Peixian
Shen, Yunhang
Qin, Yulei
Zhang, Mengdan
Lin, Xu
Yang, Jinrui
Zheng, Xiawu
Li, Ke
Sun, Xing
Wu, Yunsheng
Ji, Rongrong
Shan, Caifeng
He, Ran
contents Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2306_13394
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Fu, Chaoyou
Chen, Peixian
Shen, Yunhang
Qin, Yulei
Zhang, Mengdan
Lin, Xu
Yang, Jinrui
Zheng, Xiawu
Li, Ke
Sun, Xing
Wu, Yunsheng
Ji, Rongrong
Shan, Caifeng
He, Ran
Computer Vision and Pattern Recognition
Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation.
title MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.13394