Dynamic Evaluation of Large Language Models by Meta Probing Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Kaijie, Wang, Jindong, Zhao, Qinlin, Xu, Ruochen, Xie, Xing
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910476000034816
author Zhu, Kaijie
Wang, Jindong
Zhao, Qinlin
Xu, Ruochen
Xie, Xing
author_facet Zhu, Kaijie
Wang, Jindong
Zhao, Qinlin
Xu, Ruochen
Xie, Xing
contents Evaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination. Existing work designed evaluation protocols using well-defined algorithms for specific tasks, which cannot be easily extended to diverse scenarios. Moreover, current evaluation benchmarks can only provide the overall benchmark results and cannot support a fine-grained and multifaceted analysis of LLMs' abilities. In this paper, we propose meta probing agents (MPA), a general dynamic evaluation protocol inspired by psychometrics to evaluate LLMs. MPA is the key component of DyVal 2, which naturally extends the previous DyVal~\citep{zhu2023dyval}. MPA designs the probing and judging agents to automatically transform an original evaluation problem into a new one following psychometric theory on three basic cognitive abilities: language understanding, problem solving, and domain knowledge. These basic abilities are also dynamically configurable, allowing multifaceted analysis. We conducted extensive evaluations using MPA and found that most LLMs achieve poorer performance, indicating room for improvement. Our multifaceted analysis demonstrated the strong correlation between the basic abilities and an implicit Matthew effect on model size, i.e., larger models possess stronger correlations of the abilities. MPA can also be used as a data augmentation approach to enhance LLMs. Code is available at: https://github.com/microsoft/promptbench.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14865
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Evaluation of Large Language Models by Meta Probing Agents
Zhu, Kaijie
Wang, Jindong
Zhao, Qinlin
Xu, Ruochen
Xie, Xing
Computation and Language
Artificial Intelligence
Machine Learning
Evaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination. Existing work designed evaluation protocols using well-defined algorithms for specific tasks, which cannot be easily extended to diverse scenarios. Moreover, current evaluation benchmarks can only provide the overall benchmark results and cannot support a fine-grained and multifaceted analysis of LLMs' abilities. In this paper, we propose meta probing agents (MPA), a general dynamic evaluation protocol inspired by psychometrics to evaluate LLMs. MPA is the key component of DyVal 2, which naturally extends the previous DyVal~\citep{zhu2023dyval}. MPA designs the probing and judging agents to automatically transform an original evaluation problem into a new one following psychometric theory on three basic cognitive abilities: language understanding, problem solving, and domain knowledge. These basic abilities are also dynamically configurable, allowing multifaceted analysis. We conducted extensive evaluations using MPA and found that most LLMs achieve poorer performance, indicating room for improvement. Our multifaceted analysis demonstrated the strong correlation between the basic abilities and an implicit Matthew effect on model size, i.e., larger models possess stronger correlations of the abilities. MPA can also be used as a data augmentation approach to enhance LLMs. Code is available at: https://github.com/microsoft/promptbench.
title Dynamic Evaluation of Large Language Models by Meta Probing Agents
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.14865