A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
Fuente:
arXiv
Saved in:
| Main Authors: | Chouldechova, Alexandra, Atalla, Chad, Barocas, Solon, Cooper, A. Feder, Corvi, Emily, Dow, P. Alex, Garcia-Gathright, Jean, Pangakis, Nicholas, Reed, Stefanie, Sheng, Emily, Vann, Dan, Vogel, Matthew, Washington, Hannah, Wallach, Hanna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025)
by: Wallach, Hanna, et al.
Published: (2025)
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)
by: Chouldechova, Alexandra, et al.
Published: (2026)
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024)
by: Dow, P. Alex, et al.
Published: (2024)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
The Legal Duty to Search for Less Discriminatory Algorithms
by: Black, Emily, et al.
Published: (2024)
by: Black, Emily, et al.
Published: (2024)
What Constitutes a Less Discriminatory Algorithm?
by: Laufer, Benjamin, et al.
Published: (2024)
by: Laufer, Benjamin, et al.
Published: (2024)
Arbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification
by: Cooper, A. Feder, et al.
Published: (2023)
by: Cooper, A. Feder, et al.
Published: (2023)
Distinguishing Task-Specific and General-Purpose AI in Regulation
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Statistical Guarantees in the Search for Less Discriminatory Algorithms
by: Hays, Chris, et al.
Published: (2025)
by: Hays, Chris, et al.
Published: (2025)
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
by: Deng, Wesley Hanwen, et al.
Published: (2024)
by: Deng, Wesley Hanwen, et al.
Published: (2024)
Remote Reference Consultations Are Here to Stay
by: Reed, Emily
Published: (2021)
by: Reed, Emily
Published: (2021)
Inclusion and Empathy Are Not Enough: Cultivating Student Belonging in the Academic Library through Compassion
by: Emily Reed
Published: (2025)
by: Emily Reed
Published: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale
by: Cooper, A. Feder
Published: (2024)
by: Cooper, A. Feder
Published: (2024)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
Library Programs and Activities: Serving the Aging Directly
by: Reed, Emily W.
Published: (1973)
by: Reed, Emily W.
Published: (1973)
AI Automatons: AI Systems Intended to Imitate Humans
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
Teaching Ethical Behavior in the Global World of Information and the New AASL Standards
by: Dow, Mirah
Published: (2008)
by: Dow, Mirah
Published: (2008)
The Files are in the Computer: On Copyright, Memorization, and Generative AI
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
Eight weeks of resistance exercise improves mood state and intestinal permeability in healthy adults: A randomized controlled trial
by: Emily Dow, et al.
Published: (2025)
by: Emily Dow, et al.
Published: (2025)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
The New York State Newspaper Project Web Site.
by: Vann, William
Published: (1997)
by: Vann, William
Published: (1997)
Sustainable Spirituality: Teaching at the Intersection of Religion and Ecology
by: Jodie Vann
Published: (2025)
by: Jodie Vann
Published: (2025)
A structured regression approach for evaluating model performance across intersectional subgroups
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
by: Akpinar, Nil-Jana, et al.
Published: (2024)
by: Akpinar, Nil-Jana, et al.
Published: (2024)
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
by: Khodak, Mikhail, et al.
Published: (2024)
by: Khodak, Mikhail, et al.
Published: (2024)
Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool
by: Cheng, Lingwei, et al.
Published: (2024)
by: Cheng, Lingwei, et al.
Published: (2024)
Successful Inter-Institutional Resource Sharing in a Niche Educational Market: Formal Collaboration without a Contract
by: Dow, Elizabeth H.
Published: (2008)
by: Dow, Elizabeth H.
Published: (2008)
Similar Items
-
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025) -
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026) -
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025) -
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024) -
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)