Measuring what Matters: Construct Validity in Large Language Model Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909891512238080 |
|---|---|
| author | Bean, Andrew M. Kearns, Ryan Othniel Romanou, Angelika Hafner, Franziska Sofia Mayne, Harry Batzner, Jan Foroutan, Negar Schmitz, Chris Korgul, Karolina Batra, Hunar Deb, Oishi Beharry, Emma Emde, Cornelius Foster, Thomas Gausen, Anna Grandury, María Han, Simeng Hofmann, Valentin Ibrahim, Lujain Kim, Hazel Kirk, Hannah Rose Lin, Fangru Liu, Gabrielle Kaili-May Luettgau, Lennart Magomere, Jabez Rystrøm, Jonathan Sotnikova, Anna Yang, Yushi Zhao, Yilun Bibi, Adel Bosselut, Antoine Clark, Ronald Cohan, Arman Foerster, Jakob Gal, Yarin Hale, Scott A. Raji, Inioluwa Deborah Summerfield, Christopher Torr, Philip H. S. Ududec, Cozmin Rocher, Luc Mahdi, Adam |
| author_facet | Bean, Andrew M. Kearns, Ryan Othniel Romanou, Angelika Hafner, Franziska Sofia Mayne, Harry Batzner, Jan Foroutan, Negar Schmitz, Chris Korgul, Karolina Batra, Hunar Deb, Oishi Beharry, Emma Emde, Cornelius Foster, Thomas Gausen, Anna Grandury, María Han, Simeng Hofmann, Valentin Ibrahim, Lujain Kim, Hazel Kirk, Hannah Rose Lin, Fangru Liu, Gabrielle Kaili-May Luettgau, Lennart Magomere, Jabez Rystrøm, Jonathan Sotnikova, Anna Yang, Yushi Zhao, Yilun Bibi, Adel Bosselut, Antoine Clark, Ronald Cohan, Arman Foerster, Jakob Gal, Yarin Hale, Scott A. Raji, Inioluwa Deborah Summerfield, Christopher Torr, Philip H. S. Ududec, Cozmin Rocher, Luc Mahdi, Adam |
| contents | Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_04703 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Measuring what Matters: Construct Validity in Large Language Model Benchmarks Bean, Andrew M. Kearns, Ryan Othniel Romanou, Angelika Hafner, Franziska Sofia Mayne, Harry Batzner, Jan Foroutan, Negar Schmitz, Chris Korgul, Karolina Batra, Hunar Deb, Oishi Beharry, Emma Emde, Cornelius Foster, Thomas Gausen, Anna Grandury, María Han, Simeng Hofmann, Valentin Ibrahim, Lujain Kim, Hazel Kirk, Hannah Rose Lin, Fangru Liu, Gabrielle Kaili-May Luettgau, Lennart Magomere, Jabez Rystrøm, Jonathan Sotnikova, Anna Yang, Yushi Zhao, Yilun Bibi, Adel Bosselut, Antoine Clark, Ronald Cohan, Arman Foerster, Jakob Gal, Yarin Hale, Scott A. Raji, Inioluwa Deborah Summerfield, Christopher Torr, Philip H. S. Ududec, Cozmin Rocher, Luc Mahdi, Adam Computation and Language Artificial Intelligence Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety' and 'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. |
| title | Measuring what Matters: Construct Validity in Large Language Model Benchmarks |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2511.04703 |