A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916956169306112 |
|---|---|
| author | Shen, Ye Wang, Junying Wen, Farong Guo, Yijin Jia, Qi Zhang, Zicheng Zhai, Guangtao |
| author_facet | Shen, Ye Wang, Junying Wen, Farong Guo, Yijin Jia, Qi Zhang, Zicheng Zhai, Guangtao |
| contents | The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations suffer from high redundancy and low efficiency. Inspired by human interview processes, we propose a multi-to-one interview paradigm for efficient MLLM evaluation. Our framework consists of (i) a two-stage interview strategy with pre-interview and formal interview phases, (ii) dynamic adjustment of interviewer weights to ensure fairness, and (iii) an adaptive mechanism for question difficulty-level chosen. Experiments on different benchmarks show that the proposed paradigm achieves significantly higher correlation with full-coverage results than random sampling, with improvements of up to 17.6% in PLCC and 16.7% in SRCC, while reducing the number of required questions. These findings demonstrate that the proposed paradigm provides a reliable and efficient alternative for large-scale MLLM benchmarking. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14886 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation Shen, Ye Wang, Junying Wen, Farong Guo, Yijin Jia, Qi Zhang, Zicheng Zhai, Guangtao Computation and Language Artificial Intelligence The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations suffer from high redundancy and low efficiency. Inspired by human interview processes, we propose a multi-to-one interview paradigm for efficient MLLM evaluation. Our framework consists of (i) a two-stage interview strategy with pre-interview and formal interview phases, (ii) dynamic adjustment of interviewer weights to ensure fairness, and (iii) an adaptive mechanism for question difficulty-level chosen. Experiments on different benchmarks show that the proposed paradigm achieves significantly higher correlation with full-coverage results than random sampling, with improvements of up to 17.6% in PLCC and 16.7% in SRCC, while reducing the number of required questions. These findings demonstrate that the proposed paradigm provides a reliable and efficient alternative for large-scale MLLM benchmarking. |
| title | A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2509.14886 |