Position: AI Evaluation Should Learn from How We Test Humans

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhuang, Yan, Liu, Qi, Pardos, Zachary A., Kyllonen, Patrick C., Zu, Jiyun, Huang, Zhenya, Wang, Shijin, Chen, Enhong
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910931849576448
author Zhuang, Yan
Liu, Qi
Pardos, Zachary A.
Kyllonen, Patrick C.
Zu, Jiyun
Huang, Zhenya
Wang, Shijin
Chen, Enhong
author_facet Zhuang, Yan
Liu, Qi
Pardos, Zachary A.
Kyllonen, Patrick C.
Zu, Jiyun
Huang, Zhenya
Wang, Shijin
Chen, Enhong
contents As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically against a gold-standard test set and report metrics averaged across all items. However, this static evaluation paradigm increasingly shows its limitations, including high evaluation costs, data contamination, and the impact of low-quality or erroneous items on evaluation reliability and efficiency. In this Position, drawing from human psychometrics, we discuss a paradigm shift from static evaluation methods to adaptive testing. This involves estimating the characteristics or value of each test item in the benchmark, and tailoring each model's evaluation instead of relying on a fixed test set. This paradigm provides robust ability estimation, uncovering the latent traits underlying a model's observed scores. This position paper analyze the current possibilities, prospects, and reasons for adopting psychometrics in AI evaluation. We argue that psychometrics, a theory originating in the 20th century for human assessment, could be a powerful solution to the challenges in today's AI evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2306_10512
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Position: AI Evaluation Should Learn from How We Test Humans
Zhuang, Yan
Liu, Qi
Pardos, Zachary A.
Kyllonen, Patrick C.
Zu, Jiyun
Huang, Zhenya
Wang, Shijin
Chen, Enhong
Computation and Language
As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically against a gold-standard test set and report metrics averaged across all items. However, this static evaluation paradigm increasingly shows its limitations, including high evaluation costs, data contamination, and the impact of low-quality or erroneous items on evaluation reliability and efficiency. In this Position, drawing from human psychometrics, we discuss a paradigm shift from static evaluation methods to adaptive testing. This involves estimating the characteristics or value of each test item in the benchmark, and tailoring each model's evaluation instead of relying on a fixed test set. This paradigm provides robust ability estimation, uncovering the latent traits underlying a model's observed scores. This position paper analyze the current possibilities, prospects, and reasons for adopting psychometrics in AI evaluation. We argue that psychometrics, a theory originating in the 20th century for human assessment, could be a powerful solution to the challenges in today's AI evaluations.
title Position: AI Evaluation Should Learn from How We Test Humans
topic Computation and Language
url https://arxiv.org/abs/2306.10512