MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929499824717824 |
|---|---|
| author | Ge, Wentao Chen, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Xie, Wenya Yan, Shuo Zhu, Chenghao Lin, Ziyue Dingjie, Song Wang, Xidong Gao, Anningzhe Zhiyi, Zhang Li, Jianquan Wan, Xiang Wang, Benyou |
| author_facet | Ge, Wentao Chen, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Xie, Wenya Yan, Shuo Zhu, Chenghao Lin, Ziyue Dingjie, Song Wang, Xidong Gao, Anningzhe Zhiyi, Zhang Li, Jianquan Wan, Xiang Wang, Benyou |
| contents | Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately addressing the nuances of creative and associative multimodal tasks. However, the open-ended and subjective nature of such tasks poses a significant challenge to the evaluation methodology, where it is difficult to define the ground-truth answers for them. To this end, in our paper, we propose a new evaluation paradigm for MLLMs, which is evaluating MLLMs with per-sample criteria using potent MLLM as the judge. To validate the feasibility and effectiveness of this paradigm, we design a benchmark, dubbed MLLM-Bench, by curating the evaluation samples across six comprehensive cognitive levels. We benchmark 21 popular MLLMs in a pairwise-comparison fashion, showing diverse performance across models. Moreover, the validity of our benchmark manifests itself in reaching 88.02% agreement with human evaluation. We contend that the proposed paradigm explores the potential of MLLMs as effective evaluation tools with the help of per-sample criteria. See online leaderboard at \url{https://mllm-bench.llmzoo.com}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_13951 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria Ge, Wentao Chen, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Xie, Wenya Yan, Shuo Zhu, Chenghao Lin, Ziyue Dingjie, Song Wang, Xidong Gao, Anningzhe Zhiyi, Zhang Li, Jianquan Wan, Xiang Wang, Benyou Computation and Language Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating queries without considering user experiences, inadequately addressing the nuances of creative and associative multimodal tasks. However, the open-ended and subjective nature of such tasks poses a significant challenge to the evaluation methodology, where it is difficult to define the ground-truth answers for them. To this end, in our paper, we propose a new evaluation paradigm for MLLMs, which is evaluating MLLMs with per-sample criteria using potent MLLM as the judge. To validate the feasibility and effectiveness of this paradigm, we design a benchmark, dubbed MLLM-Bench, by curating the evaluation samples across six comprehensive cognitive levels. We benchmark 21 popular MLLMs in a pairwise-comparison fashion, showing diverse performance across models. Moreover, the validity of our benchmark manifests itself in reaching 88.02% agreement with human evaluation. We contend that the proposed paradigm explores the potential of MLLMs as effective evaluation tools with the help of per-sample criteria. See online leaderboard at \url{https://mllm-bench.llmzoo.com}. |
| title | MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2311.13951 |