Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Wallach, Hanna, Desai, Meera, Cooper, A. Feder, Wang, Angelina, Atalla, Chad, Barocas, Solon, Blodgett, Su Lin, Chouldechova, Alexandra, Corvi, Emily, Dow, P. Alex, Garcia-Gathright, Jean, Olteanu, Alexandra, Pangakis, Nicholas, Reed, Stefanie, Sheng, Emily, Vann, Dan, Vaughan, Jennifer Wortman, Vogel, Matthew, Washington, Hannah, Jacobs, Abigail Z. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024)
by: Chouldechova, Alexandra, et al.
Published: (2024)
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024)
by: Dow, P. Alex, et al.
Published: (2024)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)
by: Chouldechova, Alexandra, et al.
Published: (2026)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
by: Deng, Wesley Hanwen, et al.
Published: (2024)
by: Deng, Wesley Hanwen, et al.
Published: (2024)
AI Automatons: AI Systems Intended to Imitate Humans
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
The Legal Duty to Search for Less Discriminatory Algorithms
by: Black, Emily, et al.
Published: (2024)
by: Black, Emily, et al.
Published: (2024)
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
by: DeVrio, Alicia, et al.
Published: (2025)
by: DeVrio, Alicia, et al.
Published: (2025)
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
by: Rismani, Shalaleh, et al.
Published: (2026)
by: Rismani, Shalaleh, et al.
Published: (2026)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
What Constitutes a Less Discriminatory Algorithm?
by: Laufer, Benjamin, et al.
Published: (2024)
by: Laufer, Benjamin, et al.
Published: (2024)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
by: Hwang, Angel Hsing-Chi, et al.
Published: (2024)
by: Hwang, Angel Hsing-Chi, et al.
Published: (2024)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
Distinguishing Task-Specific and General-Purpose AI in Regulation
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Statistical Guarantees in the Search for Less Discriminatory Algorithms
by: Hays, Chris, et al.
Published: (2025)
by: Hays, Chris, et al.
Published: (2025)
ECBD: Evidence-Centered Benchmark Design for NLP
by: Liu, Yu Lu, et al.
Published: (2024)
by: Liu, Yu Lu, et al.
Published: (2024)
Remote Reference Consultations Are Here to Stay
by: Reed, Emily
Published: (2021)
by: Reed, Emily
Published: (2021)
Inclusion and Empathy Are Not Enough: Cultivating Student Belonging in the Academic Library through Compassion
by: Emily Reed
Published: (2025)
by: Emily Reed
Published: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
A structured regression approach for evaluating model performance across intersectional subgroups
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
by: Akpinar, Nil-Jana, et al.
Published: (2024)
by: Akpinar, Nil-Jana, et al.
Published: (2024)
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
by: Khodak, Mikhail, et al.
Published: (2024)
by: Khodak, Mikhail, et al.
Published: (2024)
Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool
by: Cheng, Lingwei, et al.
Published: (2024)
by: Cheng, Lingwei, et al.
Published: (2024)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
by: Pangakis, Nicholas, et al.
Published: (2024)
by: Pangakis, Nicholas, et al.
Published: (2024)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
Library Programs and Activities: Serving the Aging Directly
by: Reed, Emily W.
Published: (1973)
by: Reed, Emily W.
Published: (1973)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
by: De-Arteaga, Maria, et al.
Published: (2021)
by: De-Arteaga, Maria, et al.
Published: (2021)
On the closed neighborhood ideal of the square of the path graph
by: Olteanu, Anda, et al.
Published: (2026)
by: Olteanu, Anda, et al.
Published: (2026)
Similar Items
-
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024) -
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024) -
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024) -
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025) -
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)