TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yuxuan, Li, Tianhua, Shao, Wenqi, Zhang, Kaipeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916449712340992
author Xie, Yuxuan
Li, Tianhua
Shao, Wenqi
Zhang, Kaipeng
author_facet Xie, Yuxuan
Li, Tianhua
Shao, Wenqi
Zhang, Kaipeng
contents Recently, multimodal large language models (MLLMs) have received much attention for their impressive capabilities. The evaluation of MLLMs is becoming critical to analyzing attributes of MLLMs and providing valuable insights. However, current benchmarks overlook the problem of prompt sensitivity - minor prompt variations may lead to significant performance fluctuations. Thus, inappropriate prompts may obscure the models' capabilities, underestimating the models' performance. Moreover, different models have different preferences for different prompts, and thus, using the same prompt for all models will cause evaluation bias. This paper analyzes this deficiency in existing benchmarks and further introduces a new evaluation framework named TP-Eval, which introduces a prompt customization method to reduce evaluation biases and tap models' potential. TP-Eval will rewrite the original prompts to different customized prompts for different models. In particular, we propose some well-designed modules for prompt customization tailored to the scenario of MLLM evaluation. Extensive experiments demonstrate the effectiveness of our approach to uncovering models' capabilities, and TP-Eval should benefit the community in developing more comprehensive and convincing MLLM evaluation benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
Xie, Yuxuan
Li, Tianhua
Shao, Wenqi
Zhang, Kaipeng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recently, multimodal large language models (MLLMs) have received much attention for their impressive capabilities. The evaluation of MLLMs is becoming critical to analyzing attributes of MLLMs and providing valuable insights. However, current benchmarks overlook the problem of prompt sensitivity - minor prompt variations may lead to significant performance fluctuations. Thus, inappropriate prompts may obscure the models' capabilities, underestimating the models' performance. Moreover, different models have different preferences for different prompts, and thus, using the same prompt for all models will cause evaluation bias. This paper analyzes this deficiency in existing benchmarks and further introduces a new evaluation framework named TP-Eval, which introduces a prompt customization method to reduce evaluation biases and tap models' potential. TP-Eval will rewrite the original prompts to different customized prompts for different models. In particular, we propose some well-designed modules for prompt customization tailored to the scenario of MLLM evaluation. Extensive experiments demonstrate the effectiveness of our approach to uncovering models' capabilities, and TP-Eval should benefit the community in developing more comprehensive and convincing MLLM evaluation benchmarks.
title TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.18071