Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Angelina, Bai, Xuechunzi, Barocas, Solon, Blodgett, Su Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
di: Olteanu, Alexandra, et al.
Pubblicazione: (2025)
di: Olteanu, Alexandra, et al.
Pubblicazione: (2025)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
di: Cheng, Myra, et al.
Pubblicazione: (2024)
di: Cheng, Myra, et al.
Pubblicazione: (2024)
Distinguishing Task-Specific and General-Purpose AI in Regulation
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
di: Wang, Jennifer, et al.
Pubblicazione: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
di: Tantalaki, Nicoleta, et al.
Pubblicazione: (2025)
di: Tantalaki, Nicoleta, et al.
Pubblicazione: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
di: Harvey, Emma, et al.
Pubblicazione: (2025)
di: Harvey, Emma, et al.
Pubblicazione: (2025)
Effects of Generative AI Errors on User Reliance Across Task Difficulty
di: Anthis, Jacy Reese, et al.
Pubblicazione: (2026)
di: Anthis, Jacy Reese, et al.
Pubblicazione: (2026)
AI Generated Child Sexual Abuse Material -- What's the Harm?
di: Ciardha, Caoilte Ó, et al.
Pubblicazione: (2025)
di: Ciardha, Caoilte Ó, et al.
Pubblicazione: (2025)
Evaluating Language Models for Harmful Manipulation
di: Akbulut, Canfer, et al.
Pubblicazione: (2026)
di: Akbulut, Canfer, et al.
Pubblicazione: (2026)
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
di: Sun, Lihao, et al.
Pubblicazione: (2025)
di: Sun, Lihao, et al.
Pubblicazione: (2025)
Synthetic Data Augmentation for Enhancing Harmful Algal Bloom Detection with Machine Learning
di: Huang, Tianyi
Pubblicazione: (2025)
di: Huang, Tianyi
Pubblicazione: (2025)
The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
di: Rao, Pooja S. B., et al.
Pubblicazione: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
Designing Algorithmic Delegates: The Role of Indistinguishability in Human-AI Handoff
di: Greenwood, Sophie, et al.
Pubblicazione: (2025)
di: Greenwood, Sophie, et al.
Pubblicazione: (2025)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
di: Vaccaro, Michelle, et al.
Pubblicazione: (2026)
di: Vaccaro, Michelle, et al.
Pubblicazione: (2026)
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
di: Wu, Addison J., et al.
Pubblicazione: (2025)
di: Wu, Addison J., et al.
Pubblicazione: (2025)
Harm Amplification in Text-to-Image Models
di: Hao, Susan, et al.
Pubblicazione: (2024)
di: Hao, Susan, et al.
Pubblicazione: (2024)
Unbounded Harms, Bounded Law: Liability in the Age of Borderless AI
di: Tran, Ha-Chi
Pubblicazione: (2026)
di: Tran, Ha-Chi
Pubblicazione: (2026)
AI Automatons: AI Systems Intended to Imitate Humans
di: Olteanu, Alexandra, et al.
Pubblicazione: (2025)
di: Olteanu, Alexandra, et al.
Pubblicazione: (2025)
Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature Review
di: Pang, Rock Yuren, et al.
Pubblicazione: (2025)
di: Pang, Rock Yuren, et al.
Pubblicazione: (2025)
Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems
di: Harvey, Emma, et al.
Pubblicazione: (2024)
di: Harvey, Emma, et al.
Pubblicazione: (2024)
Harmful Suicide Content Detection
di: Park, Kyumin, et al.
Pubblicazione: (2024)
di: Park, Kyumin, et al.
Pubblicazione: (2024)
Measuring Implicit Bias in Explicitly Unbiased Large Language Models
di: Bai, Xuechunzi, et al.
Pubblicazione: (2024)
di: Bai, Xuechunzi, et al.
Pubblicazione: (2024)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
di: Jo, Claire Wonjeong, et al.
Pubblicazione: (2024)
di: Jo, Claire Wonjeong, et al.
Pubblicazione: (2024)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
di: El-Sayed, Seliem, et al.
Pubblicazione: (2024)
di: El-Sayed, Seliem, et al.
Pubblicazione: (2024)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
di: Xing, Sihao, et al.
Pubblicazione: (2026)
di: Xing, Sihao, et al.
Pubblicazione: (2026)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
di: Gringras, David
Pubblicazione: (2026)
di: Gringras, David
Pubblicazione: (2026)
What Constitutes a Less Discriminatory Algorithm?
di: Laufer, Benjamin, et al.
Pubblicazione: (2024)
di: Laufer, Benjamin, et al.
Pubblicazione: (2024)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
di: Sun, Shuzhou, et al.
Pubblicazione: (2025)
di: Sun, Shuzhou, et al.
Pubblicazione: (2025)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
di: Mohamadi, Alireza, et al.
Pubblicazione: (2025)
di: Mohamadi, Alireza, et al.
Pubblicazione: (2025)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
di: Dimino, Fabrizio, et al.
Pubblicazione: (2026)
di: Dimino, Fabrizio, et al.
Pubblicazione: (2026)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
di: Mun, Jimin, et al.
Pubblicazione: (2024)
di: Mun, Jimin, et al.
Pubblicazione: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Harm in AI-Driven Societies: An Audit of Toxicity Adoption on Chirper.ai
di: Coppolillo, Erica, et al.
Pubblicazione: (2026)
di: Coppolillo, Erica, et al.
Pubblicazione: (2026)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
di: Shieh, Evan, et al.
Pubblicazione: (2024)
di: Shieh, Evan, et al.
Pubblicazione: (2024)
Rainbow Noise: Stress-Testing Multimodal Harmful-Meme Detectors on LGBTQ Content
di: Tong, Ran, et al.
Pubblicazione: (2025)
di: Tong, Ran, et al.
Pubblicazione: (2025)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
di: Lehmann, Matthias, et al.
Pubblicazione: (2024)
di: Lehmann, Matthias, et al.
Pubblicazione: (2024)
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
di: Tahaei, Mohammmad, et al.
Pubblicazione: (2024)
di: Tahaei, Mohammmad, et al.
Pubblicazione: (2024)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
di: Cheng, Myra, et al.
Pubblicazione: (2026)
di: Cheng, Myra, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
di: Olteanu, Alexandra, et al.
Pubblicazione: (2025) -
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
di: Cheng, Myra, et al.
Pubblicazione: (2024) -
Distinguishing Task-Specific and General-Purpose AI in Regulation
di: Wang, Jennifer, et al.
Pubblicazione: (2025) -
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026) -
Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
di: Tantalaki, Nicoleta, et al.
Pubblicazione: (2025)