Dimensions of Generative AI Evaluation Design
Fuente:
arXiv
Saved in:
| Main Authors: | Dow, P. Alex, Vaughan, Jennifer Wortman, Barocas, Solon, Atalla, Chad, Chouldechova, Alexandra, Wallach, Hanna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025)
by: Wallach, Hanna, et al.
Published: (2025)
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024)
by: Chouldechova, Alexandra, et al.
Published: (2024)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
Supporting Industry Computing Researchers in Assessing, Articulating, and Addressing the Potential Negative Societal Impact of Their Work
by: Deng, Wesley Hanwen, et al.
Published: (2024)
by: Deng, Wesley Hanwen, et al.
Published: (2024)
Distinguishing Task-Specific and General-Purpose AI in Regulation
by: Wang, Jennifer, et al.
Published: (2025)
by: Wang, Jennifer, et al.
Published: (2025)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)
by: Chouldechova, Alexandra, et al.
Published: (2026)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
by: Greenwood, Sophie, et al.
Published: (2025)
by: Greenwood, Sophie, et al.
Published: (2025)
AI Automatons: AI Systems Intended to Imitate Humans
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
What Constitutes a Less Discriminatory Algorithm?
by: Laufer, Benjamin, et al.
Published: (2024)
by: Laufer, Benjamin, et al.
Published: (2024)
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics
by: Johnson, Nari, et al.
Published: (2026)
by: Johnson, Nari, et al.
Published: (2026)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
(De)Noise: Moderating the Inconsistency Between Human Decision-Makers
by: Grgić-Hlača, Nina, et al.
Published: (2024)
by: Grgić-Hlača, Nina, et al.
Published: (2024)
Statistical Guarantees in the Search for Less Discriminatory Algorithms
by: Hays, Chris, et al.
Published: (2025)
by: Hays, Chris, et al.
Published: (2025)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
by: Cooper, A. Feder, et al.
Published: (2024)
by: Cooper, A. Feder, et al.
Published: (2024)
The Legal Duty to Search for Less Discriminatory Algorithms
by: Black, Emily, et al.
Published: (2024)
by: Black, Emily, et al.
Published: (2024)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
by: Akpinar, Nil-Jana, et al.
Published: (2024)
by: Akpinar, Nil-Jana, et al.
Published: (2024)
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
by: Kapania, Shivani, et al.
Published: (2025)
by: Kapania, Shivani, et al.
Published: (2025)
A structured regression approach for evaluating model performance across intersectional subgroups
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
Generative AI for Multiple Choice STEM Assessments
by: Perdikoulias, Christina, et al.
Published: (2025)
by: Perdikoulias, Christina, et al.
Published: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
by: Pang, Rock Yuren, et al.
Published: (2025)
by: Pang, Rock Yuren, et al.
Published: (2025)
Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions
by: Vasconcelos, Helena, et al.
Published: (2023)
by: Vasconcelos, Helena, et al.
Published: (2023)
Episodic memory in AI agents poses risks that should be studied and mitigated
by: DeChant, Chad
Published: (2025)
by: DeChant, Chad
Published: (2025)
Crafting Tomorrow's Evaluations: Assessment Design Strategies in the Era of Generative AI
by: Kadel, Rajan, et al.
Published: (2024)
by: Kadel, Rajan, et al.
Published: (2024)
LLM-Augmented and Fair Machine Learning Framework for University Admission Prediction
by: Abbadi, Mohammad, et al.
Published: (2025)
by: Abbadi, Mohammad, et al.
Published: (2025)
The implicated scientist: on the role of AI researchers in the development of weapons systems
by: Volokhova, Alexandra, et al.
Published: (2026)
by: Volokhova, Alexandra, et al.
Published: (2026)
Arbitrariness and Social Prediction: The Confounding Role of Variance in Fair Classification
by: Cooper, A. Feder, et al.
Published: (2023)
by: Cooper, A. Feder, et al.
Published: (2023)
The Value of Disagreement in AI Design, Evaluation, and Alignment
by: Fazelpour, Sina, et al.
Published: (2025)
by: Fazelpour, Sina, et al.
Published: (2025)
An HCI Perspective on Sustainable GenAI Integration in Architectural Design Education
by: Nguyen, Alex Binh Vinh Duc
Published: (2026)
by: Nguyen, Alex Binh Vinh Duc
Published: (2026)
Irresponsible AI: big tech's influence on AI research and associated impacts
by: Hernandez-Garcia, Alex, et al.
Published: (2025)
by: Hernandez-Garcia, Alex, et al.
Published: (2025)
A Practical Multilevel Governance Framework for Autonomous and Intelligent Systems
by: Pöhler, Lukas D., et al.
Published: (2024)
by: Pöhler, Lukas D., et al.
Published: (2024)
Disclosure and Evaluation as Fairness Interventions for General-Purpose AI
by: Raman, Vyoma, et al.
Published: (2025)
by: Raman, Vyoma, et al.
Published: (2025)
Similar Items
-
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026) -
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024) -
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024) -
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025) -
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024)