_version_ 1866909891512238080
author Bean, Andrew M.
Kearns, Ryan Othniel
Romanou, Angelika
Hafner, Franziska Sofia
Mayne, Harry
Batzner, Jan
Foroutan, Negar
Schmitz, Chris
Korgul, Karolina
Batra, Hunar
Deb, Oishi
Beharry, Emma
Emde, Cornelius
Foster, Thomas
Gausen, Anna
Grandury, María
Han, Simeng
Hofmann, Valentin
Ibrahim, Lujain
Kim, Hazel
Kirk, Hannah Rose
Lin, Fangru
Liu, Gabrielle Kaili-May
Luettgau, Lennart
Magomere, Jabez
Rystrøm, Jonathan
Sotnikova, Anna
Yang, Yushi
Zhao, Yilun
Bibi, Adel
Bosselut, Antoine
Clark, Ronald
Cohan, Arman
Foerster, Jakob
Gal, Yarin
Hale, Scott A.
Raji, Inioluwa Deborah
Summerfield, Christopher
Torr, Philip H. S.
Ududec, Cozmin
Rocher, Luc
Mahdi, Adam
author_facet Bean, Andrew M.
Kearns, Ryan Othniel
Romanou, Angelika
Hafner, Franziska Sofia
Mayne, Harry
Batzner, Jan
Foroutan, Negar
Schmitz, Chris
Korgul, Karolina
Batra, Hunar
Deb, Oishi
Beharry, Emma
Emde, Cornelius
Foster, Thomas
Gausen, Anna
Grandury, María
Han, Simeng
Hofmann, Valentin
Ibrahim, Lujain
Kim, Hazel
Kirk, Hannah Rose
Lin, Fangru
Liu, Gabrielle Kaili-May
Luettgau, Lennart
Magomere, Jabez
Rystrøm, Jonathan
Sotnikova, Anna
Yang, Yushi
Zhao, Yilun
Bibi, Adel
Bosselut, Antoine
Clark, Ronald
Cohan, Arman
Foerster, Jakob
Gal, Yarin
Hale, Scott A.
Raji, Inioluwa Deborah
Summerfield, Christopher
Torr, Philip H. S.
Ududec, Cozmin
Rocher, Luc
Mahdi, Adam
contents Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04703
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Bean, Andrew M.
Kearns, Ryan Othniel
Romanou, Angelika
Hafner, Franziska Sofia
Mayne, Harry
Batzner, Jan
Foroutan, Negar
Schmitz, Chris
Korgul, Karolina
Batra, Hunar
Deb, Oishi
Beharry, Emma
Emde, Cornelius
Foster, Thomas
Gausen, Anna
Grandury, María
Han, Simeng
Hofmann, Valentin
Ibrahim, Lujain
Kim, Hazel
Kirk, Hannah Rose
Lin, Fangru
Liu, Gabrielle Kaili-May
Luettgau, Lennart
Magomere, Jabez
Rystrøm, Jonathan
Sotnikova, Anna
Yang, Yushi
Zhao, Yilun
Bibi, Adel
Bosselut, Antoine
Clark, Ronald
Cohan, Arman
Foerster, Jakob
Gal, Yarin
Hale, Scott A.
Raji, Inioluwa Deborah
Summerfield, Christopher
Torr, Philip H. S.
Ududec, Cozmin
Rocher, Luc
Mahdi, Adam
Computation and Language
Artificial Intelligence
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.
title Measuring what Matters: Construct Validity in Large Language Model Benchmarks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.04703