Evaluating Generative AI Systems is a Social Science Measurement Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wallach, Hanna, Desai, Meera, Pangakis, Nicholas, Cooper, A. Feder, Wang, Angelina, Barocas, Solon, Chouldechova, Alexandra, Atalla, Chad, Blodgett, Su Lin, Corvi, Emily, Dow, P. Alex, Garcia-Gathright, Jean, Olteanu, Alexandra, Reed, Stefanie, Sheng, Emily, Vann, Dan, Vaughan, Jennifer Wortman, Vogel, Matthew, Washington, Hannah, Jacobs, Abigail Z.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916485623971840
author Wallach, Hanna
Desai, Meera
Pangakis, Nicholas
Cooper, A. Feder
Wang, Angelina
Barocas, Solon
Chouldechova, Alexandra
Atalla, Chad
Blodgett, Su Lin
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Olteanu, Alexandra
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vaughan, Jennifer Wortman
Vogel, Matthew
Washington, Hannah
Jacobs, Abigail Z.
author_facet Wallach, Hanna
Desai, Meera
Pangakis, Nicholas
Cooper, A. Feder
Wang, Angelina
Barocas, Solon
Chouldechova, Alexandra
Atalla, Chad
Blodgett, Su Lin
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Olteanu, Alexandra
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vaughan, Jennifer Wortman
Vogel, Matthew
Washington, Hannah
Jacobs, Abigail Z.
contents Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that these measurement tasks are highly reminiscent of measurement tasks found throughout the social sciences. With this in mind, we present a framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, impacts, opportunities, and risks of GenAI systems. The framework distinguishes between four levels: the background concept, the systematized concept, the measurement instrument(s), and the instance-level measurements themselves. This four-level approach differs from the way measurement is typically done in ML, where researchers and practitioners appear to jump straight from background concepts to measurement instruments, with little to no explicit systematization in between. As well as surfacing assumptions, thereby making it easier to understand exactly what the resulting measurements do and do not mean, this framework has two important implications for evaluating evaluations: First, it can enable stakeholders from different worlds to participate in conceptual debates, broadening the expertise involved in evaluating GenAI systems. Second, it brings rigor to operational debates by offering a set of lenses for interrogating the validity of measurement instruments and their resulting measurements.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Generative AI Systems is a Social Science Measurement Challenge
Wallach, Hanna
Desai, Meera
Pangakis, Nicholas
Cooper, A. Feder
Wang, Angelina
Barocas, Solon
Chouldechova, Alexandra
Atalla, Chad
Blodgett, Su Lin
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Olteanu, Alexandra
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vaughan, Jennifer Wortman
Vogel, Matthew
Washington, Hannah
Jacobs, Abigail Z.
Computers and Society
Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that these measurement tasks are highly reminiscent of measurement tasks found throughout the social sciences. With this in mind, we present a framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, impacts, opportunities, and risks of GenAI systems. The framework distinguishes between four levels: the background concept, the systematized concept, the measurement instrument(s), and the instance-level measurements themselves. This four-level approach differs from the way measurement is typically done in ML, where researchers and practitioners appear to jump straight from background concepts to measurement instruments, with little to no explicit systematization in between. As well as surfacing assumptions, thereby making it easier to understand exactly what the resulting measurements do and do not mean, this framework has two important implications for evaluating evaluations: First, it can enable stakeholders from different worlds to participate in conceptual debates, broadening the expertise involved in evaluating GenAI systems. Second, it brings rigor to operational debates by offering a set of lenses for interrogating the validity of measurement instruments and their resulting measurements.
title Evaluating Generative AI Systems is a Social Science Measurement Challenge
topic Computers and Society
url https://arxiv.org/abs/2411.10939