Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Harvey, Emma, Sheng, Emily, Blodgett, Su Lin, Chouldechova, Alexandra, Garcia-Gathright, Jean, Olteanu, Alexandra, Wallach, Hanna |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024)
by: Harvey, Emma, et al.
Published: (2024)
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025)
by: Corvi, Emily, et al.
Published: (2025)
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023)
by: Lucy, Li, et al.
Published: (2023)
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026)
by: Agarwal, Dhruv, et al.
Published: (2026)
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)
by: Guerdan, Luke, et al.
Published: (2024)
Evaluating Generative AI Systems is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2024)
by: Wallach, Hanna, et al.
Published: (2024)
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
by: Wallach, Hanna, et al.
Published: (2025)
by: Wallach, Hanna, et al.
Published: (2025)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
by: Guerdan, Luke, et al.
Published: (2025)
by: Guerdan, Luke, et al.
Published: (2025)
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
by: Chouldechova, Alexandra, et al.
Published: (2024)
by: Chouldechova, Alexandra, et al.
Published: (2024)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language Technologies
by: DeVrio, Alicia, et al.
Published: (2025)
by: DeVrio, Alicia, et al.
Published: (2025)
Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
From Use to Oversight: How Mental Models Influence User Behavior and Output in AI Writing Assistants
by: Rismani, Shalaleh, et al.
Published: (2026)
by: Rismani, Shalaleh, et al.
Published: (2026)
"It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
by: Hwang, Angel Hsing-Chi, et al.
Published: (2024)
by: Hwang, Angel Hsing-Chi, et al.
Published: (2024)
Dimensions of Generative AI Evaluation Design
by: Dow, P. Alex, et al.
Published: (2024)
by: Dow, P. Alex, et al.
Published: (2024)
AI Automatons: AI Systems Intended to Imitate Humans
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
ECBD: Evidence-Centered Benchmark Design for NLP
by: Liu, Yu Lu, et al.
Published: (2024)
by: Liu, Yu Lu, et al.
Published: (2024)
Do Responsible AI Artifacts Advance Stakeholder Goals? Four Key Barriers Perceived by Legal and Civil Stakeholders
by: Kawakami, Anna, et al.
Published: (2024)
by: Kawakami, Anna, et al.
Published: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
"Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
by: Akpinar, Nil-Jana, et al.
Published: (2024)
by: Akpinar, Nil-Jana, et al.
Published: (2024)
A structured regression approach for evaluating model performance across intersectional subgroups
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
Challenges to Evaluating the Generalization of Coreference Resolution Models: A Measurement Modeling Perspective
by: Porada, Ian, et al.
Published: (2023)
by: Porada, Ian, et al.
Published: (2023)
Algorithm-Assisted Decision Making and Racial Disparities in Housing: A Study of the Allegheny Housing Assessment Tool
by: Cheng, Lingwei, et al.
Published: (2024)
by: Cheng, Lingwei, et al.
Published: (2024)
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
by: Olteanu, Alexandra, et al.
Published: (2025)
by: Olteanu, Alexandra, et al.
Published: (2025)
A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
by: Harvey, Emma, et al.
Published: (2025)
by: Harvey, Emma, et al.
Published: (2025)
Leveraging Expert Consistency to Improve Algorithmic Decision Support
by: De-Arteaga, Maria, et al.
Published: (2021)
by: De-Arteaga, Maria, et al.
Published: (2021)
Understanding the Dataset Practitioners Behind Large Language Model Development
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
by: Chouldechova, Alexandra, et al.
Published: (2026)
by: Chouldechova, Alexandra, et al.
Published: (2026)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
by: Lehmann, Matthias, et al.
Published: (2024)
by: Lehmann, Matthias, et al.
Published: (2024)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
by: Fleisig, Eve, et al.
Published: (2024)
by: Fleisig, Eve, et al.
Published: (2024)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
by: Anthis, Jacy Reese, et al.
Published: (2026)
by: Anthis, Jacy Reese, et al.
Published: (2026)
"It Was a Magical Box": Understanding Practitioner Workflows and Needs in Optimization
by: Lawless, Connor, et al.
Published: (2025)
by: Lawless, Connor, et al.
Published: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
by: Chien, Jennifer, et al.
Published: (2024)
by: Chien, Jennifer, et al.
Published: (2024)
Fairness-in-the-Workflow: How Machine Learning Practitioners at Big Tech Companies Approach Fairness in Recommender Systems
by: Yan, Jing Nathan, et al.
Published: (2025)
by: Yan, Jing Nathan, et al.
Published: (2025)
When Large Language Model Meets Optimization
by: Huang, Sen, et al.
Published: (2024)
by: Huang, Sen, et al.
Published: (2024)
Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities
by: Nguyen, Ilana, et al.
Published: (2026)
by: Nguyen, Ilana, et al.
Published: (2026)
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification
by: Shcharbakova, Hanna, et al.
Published: (2025)
by: Shcharbakova, Hanna, et al.
Published: (2025)
Similar Items
-
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
by: Harvey, Emma, et al.
Published: (2024) -
Taxonomizing Representational Harms using Speech Act Theory
by: Corvi, Emily, et al.
Published: (2025) -
"One-Size-Fits-All"? Examining Expectations around What Constitute "Fair" or "Good" NLG System Behaviors
by: Lucy, Li, et al.
Published: (2023) -
AI-Assisted Systematization for Evaluating GenAI Systems
by: Agarwal, Dhruv, et al.
Published: (2026) -
A Framework for Evaluating LLMs Under Task Indeterminacy
by: Guerdan, Luke, et al.
Published: (2024)