A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chouldechova, Alexandra, Atalla, Chad, Barocas, Solon, Cooper, A. Feder, Corvi, Emily, Dow, P. Alex, Garcia-Gathright, Jean, Pangakis, Nicholas, Reed, Stefanie, Sheng, Emily, Vann, Dan, Vogel, Matthew, Washington, Hannah, Wallach, Hanna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912141632602112
author Chouldechova, Alexandra
Atalla, Chad
Barocas, Solon
Cooper, A. Feder
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Pangakis, Nicholas
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vogel, Matthew
Washington, Hannah
Wallach, Hanna
author_facet Chouldechova, Alexandra
Atalla, Chad
Barocas, Solon
Cooper, A. Feder
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Pangakis, Nicholas
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vogel, Matthew
Washington, Hannah
Wallach, Hanna
contents The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard for valid measurement that helps place many of the disparate-seeming evaluation practices in use today on a common footing. Our framework, grounded in measurement theory from the social sciences, extends the work of Adcock & Collier (2001) in which the authors formalized valid measurement of concepts in political science via three processes: systematizing background concepts, operationalizing systematized concepts via annotation procedures, and applying those procedures to instances. We argue that valid measurement of GenAI systems' capabilities, risks, and impacts, further requires systematizing, operationalizing, and applying not only the entailed concepts, but also the contexts of interest and the metrics used. This involves both descriptive reasoning about particular instances and inferential reasoning about underlying populations, which is the purview of statistics. By placing many disparate-seeming GenAI evaluation practices on a common footing, our framework enables individual evaluations to be better understood, interrogated for reliability and validity, and meaningfully compared. This is an important step in advancing GenAI evaluation practices toward more formalized and theoretically grounded processes -- i.e., toward a science of GenAI evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01934
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
Chouldechova, Alexandra
Atalla, Chad
Barocas, Solon
Cooper, A. Feder
Corvi, Emily
Dow, P. Alex
Garcia-Gathright, Jean
Pangakis, Nicholas
Reed, Stefanie
Sheng, Emily
Vann, Dan
Vogel, Matthew
Washington, Hannah
Wallach, Hanna
Computers and Society
The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard for valid measurement that helps place many of the disparate-seeming evaluation practices in use today on a common footing. Our framework, grounded in measurement theory from the social sciences, extends the work of Adcock & Collier (2001) in which the authors formalized valid measurement of concepts in political science via three processes: systematizing background concepts, operationalizing systematized concepts via annotation procedures, and applying those procedures to instances. We argue that valid measurement of GenAI systems' capabilities, risks, and impacts, further requires systematizing, operationalizing, and applying not only the entailed concepts, but also the contexts of interest and the metrics used. This involves both descriptive reasoning about particular instances and inferential reasoning about underlying populations, which is the purview of statistics. By placing many disparate-seeming GenAI evaluation practices on a common footing, our framework enables individual evaluations to be better understood, interrogated for reliability and validity, and meaningfully compared. This is an important step in advancing GenAI evaluation practices toward more formalized and theoretically grounded processes -- i.e., toward a science of GenAI evaluations.
title A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
topic Computers and Society
url https://arxiv.org/abs/2412.01934