MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ge, Wentao, Chen, Shunian, Chen, Guiming Hardy, Chen, Junying, Chen, Zhihong, Chen, Nuo, Xie, Wenya, Yan, Shuo, Zhu, Chenghao, Lin, Ziyue, Dingjie, Song, Wang, Xidong, Gao, Anningzhe, Zhiyi, Zhang, Li, Jianquan, Wan, Xiang, Wang, Benyou
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929499824717824
author Ge, Wentao
Chen, Shunian
Chen, Guiming Hardy
Chen, Junying
Chen, Zhihong
Chen, Nuo
Xie, Wenya
Yan, Shuo
Zhu, Chenghao
Lin, Ziyue
Dingjie, Song
Wang, Xidong
Gao, Anningzhe
Zhiyi, Zhang
Li, Jianquan
Wan, Xiang
Wang, Benyou
author_facet Ge, Wentao
Chen, Shunian
Chen, Guiming Hardy
Chen, Junying
Chen, Zhihong
Chen, Nuo
Xie, Wenya
Yan, Shuo
Zhu, Chenghao
Lin, Ziyue
Dingjie, Song
Wang, Xidong
Gao, Anningzhe
Zhiyi, Zhang
Li, Jianquan
Wan, Xiang
Wang, Benyou
contents Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately addressing the nuances of creative and associative multimodal tasks. However, the open-ended and subjective nature of such tasks poses a significant challenge to the evaluation methodology, where it is difficult to define the ground-truth answers for them. To this end, in our paper, we propose a new evaluation paradigm for MLLMs, which is evaluating MLLMs with per-sample criteria using potent MLLM as the judge. To validate the feasibility and effectiveness of this paradigm, we design a benchmark, dubbed MLLM-Bench, by curating the evaluation samples across six comprehensive cognitive levels. We benchmark 21 popular MLLMs in a pairwise-comparison fashion, showing diverse performance across models. Moreover, the validity of our benchmark manifests itself in reaching 88.02% agreement with human evaluation. We contend that the proposed paradigm explores the potential of MLLMs as effective evaluation tools with the help of per-sample criteria. See online leaderboard at \url{https://mllm-bench.llmzoo.com}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_13951
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
Ge, Wentao
Chen, Shunian
Chen, Guiming Hardy
Chen, Junying
Chen, Zhihong
Chen, Nuo
Xie, Wenya
Yan, Shuo
Zhu, Chenghao
Lin, Ziyue
Dingjie, Song
Wang, Xidong
Gao, Anningzhe
Zhiyi, Zhang
Li, Jianquan
Wan, Xiang
Wang, Benyou
Computation and Language
Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately addressing the nuances of creative and associative multimodal tasks. However, the open-ended and subjective nature of such tasks poses a significant challenge to the evaluation methodology, where it is difficult to define the ground-truth answers for them. To this end, in our paper, we propose a new evaluation paradigm for MLLMs, which is evaluating MLLMs with per-sample criteria using potent MLLM as the judge. To validate the feasibility and effectiveness of this paradigm, we design a benchmark, dubbed MLLM-Bench, by curating the evaluation samples across six comprehensive cognitive levels. We benchmark 21 popular MLLMs in a pairwise-comparison fashion, showing diverse performance across models. Moreover, the validity of our benchmark manifests itself in reaching 88.02% agreement with human evaluation. We contend that the proposed paradigm explores the potential of MLLMs as effective evaluation tools with the help of per-sample criteria. See online leaderboard at \url{https://mllm-bench.llmzoo.com}.
title MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
topic Computation and Language
url https://arxiv.org/abs/2311.13951