Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
Fuente:
arXiv
Saved in:
| Main Authors: | Tantalaki, Nicoleta, Vei, Sophia, Vakali, Athena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rolling in the deep of cognitive and AI biases
by: Tantalaki, Nicoleta, et al.
Published: (2024)
by: Tantalaki, Nicoleta, et al.
Published: (2024)
FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines
by: Vakali, Athena, et al.
Published: (2025)
by: Vakali, Athena, et al.
Published: (2025)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
by: Mun, Jimin, et al.
Published: (2024)
by: Mun, Jimin, et al.
Published: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
AI Harmonics: a human-centric and harms severity-adaptive AI risk assessment framework
by: Vei, Sofia, et al.
Published: (2025)
by: Vei, Sofia, et al.
Published: (2025)
Harm in AI-Driven Societies: An Audit of Toxicity Adoption on Chirper.ai
by: Coppolillo, Erica, et al.
Published: (2026)
by: Coppolillo, Erica, et al.
Published: (2026)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
by: Xing, Sihao, et al.
Published: (2026)
by: Xing, Sihao, et al.
Published: (2026)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
by: Sun, Shuzhou, et al.
Published: (2025)
by: Sun, Shuzhou, et al.
Published: (2025)
Unbounded Harms, Bounded Law: Liability in the Age of Borderless AI
by: Tran, Ha-Chi
Published: (2026)
by: Tran, Ha-Chi
Published: (2026)
A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms
by: Abercrombie, Gavin, et al.
Published: (2024)
by: Abercrombie, Gavin, et al.
Published: (2024)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
by: El-Sayed, Seliem, et al.
Published: (2024)
by: El-Sayed, Seliem, et al.
Published: (2024)
Evaluating Language Models for Harmful Manipulation
by: Akbulut, Canfer, et al.
Published: (2026)
by: Akbulut, Canfer, et al.
Published: (2026)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
by: Wang, Angelina, et al.
Published: (2024)
by: Wang, Angelina, et al.
Published: (2024)
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
by: Ehsan, Upol, et al.
Published: (2026)
by: Ehsan, Upol, et al.
Published: (2026)
AI Generated Child Sexual Abuse Material -- What's the Harm?
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
by: Ciardha, Caoilte Ó, et al.
Published: (2025)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
by: Rao, Pooja S. B., et al.
Published: (2025)
by: Rao, Pooja S. B., et al.
Published: (2025)
Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
by: Paraschou, Eva, et al.
Published: (2025)
by: Paraschou, Eva, et al.
Published: (2025)
Harm Amplification in Text-to-Image Models
by: Hao, Susan, et al.
Published: (2024)
by: Hao, Susan, et al.
Published: (2024)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
by: Jo, Claire Wonjeong, et al.
Published: (2024)
by: Jo, Claire Wonjeong, et al.
Published: (2024)
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
by: Tahaei, Mohammmad, et al.
Published: (2024)
by: Tahaei, Mohammmad, et al.
Published: (2024)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
by: Gringras, David
Published: (2026)
by: Gringras, David
Published: (2026)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
by: Lehmann, Matthias, et al.
Published: (2024)
by: Lehmann, Matthias, et al.
Published: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
by: Chien, Jennifer, et al.
Published: (2024)
by: Chien, Jennifer, et al.
Published: (2024)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
by: Drinkall, Toby
Published: (2025)
by: Drinkall, Toby
Published: (2025)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
by: Dimino, Fabrizio, et al.
Published: (2026)
by: Dimino, Fabrizio, et al.
Published: (2026)
Generative Ghosts: Anticipating Benefits and Risks of AI Afterlives
by: Morris, Meredith Ringel, et al.
Published: (2024)
by: Morris, Meredith Ringel, et al.
Published: (2024)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
by: Fan, Xianzhe, et al.
Published: (2024)
by: Fan, Xianzhe, et al.
Published: (2024)
Recognising, Anticipating, and Mitigating LLM Pollution of Online Behavioural Research
by: Rilla, Raluca, et al.
Published: (2025)
by: Rilla, Raluca, et al.
Published: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Echoes of Power: Investigating Geopolitical Bias in US and China Large Language Models
by: Pacheco, Andre G. C., et al.
Published: (2025)
by: Pacheco, Andre G. C., et al.
Published: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
by: Choi, Sooyung, et al.
Published: (2025)
by: Choi, Sooyung, et al.
Published: (2025)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
by: Shieh, Evan, et al.
Published: (2024)
by: Shieh, Evan, et al.
Published: (2024)
AI Consciousness and Public Perceptions: Four Futures
by: Fernandez, Ines, et al.
Published: (2024)
by: Fernandez, Ines, et al.
Published: (2024)
Equity Bias: An Ethical Framework for AI Design
by: Lockwood, Mary
Published: (2026)
by: Lockwood, Mary
Published: (2026)
RealHarm: A Collection of Real-World Language Model Application Failures
by: Jeune, Pierre Le, et al.
Published: (2025)
by: Jeune, Pierre Le, et al.
Published: (2025)
Similar Items
-
Rolling in the deep of cognitive and AI biases
by: Tantalaki, Nicoleta, et al.
Published: (2024) -
FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines
by: Vakali, Athena, et al.
Published: (2025) -
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
by: Mun, Jimin, et al.
Published: (2024) -
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026) -
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)