A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Ye, Wang, Junying, Wen, Farong, Guo, Yijin, Jia, Qi, Zhang, Zicheng, Zhai, Guangtao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916956169306112
author Shen, Ye
Wang, Junying
Wen, Farong
Guo, Yijin
Jia, Qi
Zhang, Zicheng
Zhai, Guangtao
author_facet Shen, Ye
Wang, Junying
Wen, Farong
Guo, Yijin
Jia, Qi
Zhang, Zicheng
Zhai, Guangtao
contents The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations suffer from high redundancy and low efficiency. Inspired by human interview processes, we propose a multi-to-one interview paradigm for efficient MLLM evaluation. Our framework consists of (i) a two-stage interview strategy with pre-interview and formal interview phases, (ii) dynamic adjustment of interviewer weights to ensure fairness, and (iii) an adaptive mechanism for question difficulty-level chosen. Experiments on different benchmarks show that the proposed paradigm achieves significantly higher correlation with full-coverage results than random sampling, with improvements of up to 17.6% in PLCC and 16.7% in SRCC, while reducing the number of required questions. These findings demonstrate that the proposed paradigm provides a reliable and efficient alternative for large-scale MLLM benchmarking.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
Shen, Ye
Wang, Junying
Wen, Farong
Guo, Yijin
Jia, Qi
Zhang, Zicheng
Zhai, Guangtao
Computation and Language
Artificial Intelligence
The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations suffer from high redundancy and low efficiency. Inspired by human interview processes, we propose a multi-to-one interview paradigm for efficient MLLM evaluation. Our framework consists of (i) a two-stage interview strategy with pre-interview and formal interview phases, (ii) dynamic adjustment of interviewer weights to ensure fairness, and (iii) an adaptive mechanism for question difficulty-level chosen. Experiments on different benchmarks show that the proposed paradigm achieves significantly higher correlation with full-coverage results than random sampling, with improvements of up to 17.6% in PLCC and 16.7% in SRCC, while reducing the number of required questions. These findings demonstrate that the proposed paradigm provides a reliable and efficient alternative for large-scale MLLM benchmarking.
title A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.14886