Evaluating Generative AI Systems is a Social Science Measurement Challenge
Fuente:
arXiv
Guardado en:
| Autores principales: | Wallach, Hanna, Desai, Meera, Pangakis, Nicholas, Cooper, A. Feder, Wang, Angelina, Barocas, Solon, Chouldechova, Alexandra, Atalla, Chad, Blodgett, Su Lin, Corvi, Emily, Dow, P. Alex, Garcia-Gathright, Jean, Olteanu, Alexandra, Reed, Stefanie, Sheng, Emily, Vann, Dan, Vaughan, Jennifer Wortman, Vogel, Matthew, Washington, Hannah, Jacobs, Abigail Z. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
por: Wallach, Hanna, et al.
Publicado: (2025)
por: Wallach, Hanna, et al.
Publicado: (2025)
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
por: Chouldechova, Alexandra, et al.
Publicado: (2024)
por: Chouldechova, Alexandra, et al.
Publicado: (2024)
Dimensions of Generative AI Evaluation Design
por: Dow, P. Alex, et al.
Publicado: (2024)
por: Dow, P. Alex, et al.
Publicado: (2024)
Taxonomizing Representational Harms using Speech Act Theory
por: Corvi, Emily, et al.
Publicado: (2025)
por: Corvi, Emily, et al.
Publicado: (2025)
AI-Assisted Systematization for Evaluating GenAI Systems
por: Agarwal, Dhruv, et al.
Publicado: (2026)
por: Agarwal, Dhruv, et al.
Publicado: (2026)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
por: Chouldechova, Alexandra, et al.
Publicado: (2026)
por: Chouldechova, Alexandra, et al.
Publicado: (2026)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
por: Harvey, Emma, et al.
Publicado: (2024)
por: Harvey, Emma, et al.
Publicado: (2024)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
por: Harvey, Emma, et al.
Publicado: (2025)
por: Harvey, Emma, et al.
Publicado: (2025)
A Framework for Evaluating LLMs Under Task Indeterminacy
por: Guerdan, Luke, et al.
Publicado: (2024)
por: Guerdan, Luke, et al.
Publicado: (2024)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
por: Guerdan, Luke, et al.
Publicado: (2025)
por: Guerdan, Luke, et al.
Publicado: (2025)
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
por: Deng, Wesley Hanwen, et al.
Publicado: (2024)
por: Deng, Wesley Hanwen, et al.
Publicado: (2024)
AI Automatons: AI Systems Intended to Imitate Humans
por: Olteanu, Alexandra, et al.
Publicado: (2025)
por: Olteanu, Alexandra, et al.
Publicado: (2025)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
por: Wang, Angelina, et al.
Publicado: (2024)
por: Wang, Angelina, et al.
Publicado: (2024)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
por: Anthis, Jacy Reese, et al.
Publicado: (2026)
por: Anthis, Jacy Reese, et al.
Publicado: (2026)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
por: Lucy, Li, et al.
Publicado: (2023)
por: Lucy, Li, et al.
Publicado: (2023)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
por: Olteanu, Alexandra, et al.
Publicado: (2025)
por: Olteanu, Alexandra, et al.
Publicado: (2025)
The Legal Duty to Search for Less Discriminatory Algorithms
por: Black, Emily, et al.
Publicado: (2024)
por: Black, Emily, et al.
Publicado: (2024)
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
por: DeVrio, Alicia, et al.
Publicado: (2025)
por: DeVrio, Alicia, et al.
Publicado: (2025)
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
por: Rismani, Shalaleh, et al.
Publicado: (2026)
por: Rismani, Shalaleh, et al.
Publicado: (2026)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
por: Cheng, Myra, et al.
Publicado: (2024)
por: Cheng, Myra, et al.
Publicado: (2024)
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
por: Cheng, Myra, et al.
Publicado: (2025)
por: Cheng, Myra, et al.
Publicado: (2025)
What Constitutes a Less Discriminatory Algorithm?
por: Laufer, Benjamin, et al.
Publicado: (2024)
por: Laufer, Benjamin, et al.
Publicado: (2024)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
por: Hwang, Angel Hsing-Chi, et al.
Publicado: (2024)
por: Hwang, Angel Hsing-Chi, et al.
Publicado: (2024)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
por: Kawakami, Anna, et al.
Publicado: (2024)
por: Kawakami, Anna, et al.
Publicado: (2024)
Distinguishing Task-Specific and General-Purpose AI in Regulation
por: Wang, Jennifer, et al.
Publicado: (2025)
por: Wang, Jennifer, et al.
Publicado: (2025)
Statistical Guarantees in the Search for Less Discriminatory Algorithms
por: Hays, Chris, et al.
Publicado: (2025)
por: Hays, Chris, et al.
Publicado: (2025)
ECBD: Evidence-Centered Benchmark Design for NLP
por: Liu, Yu Lu, et al.
Publicado: (2024)
por: Liu, Yu Lu, et al.
Publicado: (2024)
Remote Reference Consultations Are Here to Stay
por: Reed, Emily
Publicado: (2021)
por: Reed, Emily
Publicado: (2021)
Inclusion and Empathy Are Not Enough: Cultivating Student Belonging in the Academic Library through Compassion
por: Emily Reed
Publicado: (2025)
por: Emily Reed
Publicado: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
por: Pang, Rock Yuren, et al.
Publicado: (2025)
por: Pang, Rock Yuren, et al.
Publicado: (2025)
A structured regression approach for evaluating model performance across intersectional subgroups
por: Herlihy, Christine, et al.
Publicado: (2024)
por: Herlihy, Christine, et al.
Publicado: (2024)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
por: Akpinar, Nil-Jana, et al.
Publicado: (2024)
por: Akpinar, Nil-Jana, et al.
Publicado: (2024)
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
por: Khodak, Mikhail, et al.
Publicado: (2024)
por: Khodak, Mikhail, et al.
Publicado: (2024)
Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool
por: Cheng, Lingwei, et al.
Publicado: (2024)
por: Cheng, Lingwei, et al.
Publicado: (2024)
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI
por: Pangakis, Nicholas, et al.
Publicado: (2024)
por: Pangakis, Nicholas, et al.
Publicado: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
por: Pangakis, Nicholas, et al.
Publicado: (2024)
por: Pangakis, Nicholas, et al.
Publicado: (2024)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
por: Greenwood, Sophie, et al.
Publicado: (2025)
por: Greenwood, Sophie, et al.
Publicado: (2025)
Library Programs and Activities: Serving the Aging Directly
por: Reed, Emily W.
Publicado: (1973)
por: Reed, Emily W.
Publicado: (1973)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
por: De-Arteaga, Maria, et al.
Publicado: (2021)
por: De-Arteaga, Maria, et al.
Publicado: (2021)
On the closed neighborhood ideal of the square of the path graph
por: Olteanu, Anda, et al.
Publicado: (2026)
por: Olteanu, Anda, et al.
Publicado: (2026)
Ejemplares similares
-
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
por: Wallach, Hanna, et al.
Publicado: (2025) -
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
por: Chouldechova, Alexandra, et al.
Publicado: (2024) -
Dimensions of Generative AI Evaluation Design
por: Dow, P. Alex, et al.
Publicado: (2024) -
Taxonomizing Representational Harms using Speech Act Theory
por: Corvi, Emily, et al.
Publicado: (2025) -
AI-Assisted Systematization for Evaluating GenAI Systems
por: Agarwal, Dhruv, et al.
Publicado: (2026)