Echoes of AI Harms: A Human-LLM Synergistic Framework for Bias-Driven Harm Anticipation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tantalaki, Nicoleta, Vei, Sophia, Vakali, Athena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rolling in the deep of cognitive and AI biases
von: Tantalaki, Nicoleta, et al.
Veröffentlicht: (2024)
von: Tantalaki, Nicoleta, et al.
Veröffentlicht: (2024)
FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines
von: Vakali, Athena, et al.
Veröffentlicht: (2025)
von: Vakali, Athena, et al.
Veröffentlicht: (2025)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
von: Mun, Jimin, et al.
Veröffentlicht: (2024)
von: Mun, Jimin, et al.
Veröffentlicht: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)
AI Harmonics: a human-centric and harms severity-adaptive AI risk assessment framework
von: Vei, Sofia, et al.
Veröffentlicht: (2025)
von: Vei, Sofia, et al.
Veröffentlicht: (2025)
Harm in AI-Driven Societies: An Audit of Toxicity Adoption on Chirper.ai
von: Coppolillo, Erica, et al.
Veröffentlicht: (2026)
von: Coppolillo, Erica, et al.
Veröffentlicht: (2026)
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
von: Xing, Sihao, et al.
Veröffentlicht: (2026)
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
von: Sun, Shuzhou, et al.
Veröffentlicht: (2025)
Unbounded Harms, Bounded Law: Liability in the Age of Borderless AI
von: Tran, Ha-Chi
Veröffentlicht: (2026)
von: Tran, Ha-Chi
Veröffentlicht: (2026)
A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms
von: Abercrombie, Gavin, et al.
Veröffentlicht: (2024)
von: Abercrombie, Gavin, et al.
Veröffentlicht: (2024)
A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI
von: El-Sayed, Seliem, et al.
Veröffentlicht: (2024)
von: El-Sayed, Seliem, et al.
Veröffentlicht: (2024)
Evaluating Language Models for Harmful Manipulation
von: Akbulut, Canfer, et al.
Veröffentlicht: (2026)
von: Akbulut, Canfer, et al.
Veröffentlicht: (2026)
Measuring Machine Learning Harms from Stereotypes Requires Understanding Who Is Harmed by Which Errors in What Ways
von: Wang, Angelina, et al.
Veröffentlicht: (2024)
von: Wang, Angelina, et al.
Veröffentlicht: (2024)
From Future of Work to Future of Workers: Addressing Asymptomatic AI Harms for Dignified Human-AI Interaction
von: Ehsan, Upol, et al.
Veröffentlicht: (2026)
von: Ehsan, Upol, et al.
Veröffentlicht: (2026)
AI Generated Child Sexual Abuse Material -- What's the Harm?
von: Ciardha, Caoilte Ó, et al.
Veröffentlicht: (2025)
von: Ciardha, Caoilte Ó, et al.
Veröffentlicht: (2025)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
von: Mohamadi, Alireza, et al.
Veröffentlicht: (2025)
The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
von: Rao, Pooja S. B., et al.
Veröffentlicht: (2025)
von: Rao, Pooja S. B., et al.
Veröffentlicht: (2025)
Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
von: Paraschou, Eva, et al.
Veröffentlicht: (2025)
von: Paraschou, Eva, et al.
Veröffentlicht: (2025)
Harm Amplification in Text-to-Image Models
von: Hao, Susan, et al.
Veröffentlicht: (2024)
von: Hao, Susan, et al.
Veröffentlicht: (2024)
Harmful YouTube Video Detection: A Taxonomy of Online Harm and MLLMs as Alternative Annotators
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
von: Jo, Claire Wonjeong, et al.
Veröffentlicht: (2024)
Surveys Considered Harmful? Reflecting on the Use of Surveys in AI Research, Development, and Governance
von: Tahaei, Mohammmad, et al.
Veröffentlicht: (2024)
von: Tahaei, Mohammmad, et al.
Veröffentlicht: (2024)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
von: Gautam, Sanjana, et al.
Veröffentlicht: (2024)
Harmful Suicide Content Detection
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
von: Park, Kyumin, et al.
Veröffentlicht: (2024)
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
von: Gringras, David
Veröffentlicht: (2026)
von: Gringras, David
Veröffentlicht: (2026)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2024)
AI Meets the Classroom: When Do Large Language Models Harm Learning?
von: Lehmann, Matthias, et al.
Veröffentlicht: (2024)
von: Lehmann, Matthias, et al.
Veröffentlicht: (2024)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
von: Chien, Jennifer, et al.
Veröffentlicht: (2024)
von: Chien, Jennifer, et al.
Veröffentlicht: (2024)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
von: Drinkall, Toby
Veröffentlicht: (2025)
von: Drinkall, Toby
Veröffentlicht: (2025)
Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2026)
von: Dimino, Fabrizio, et al.
Veröffentlicht: (2026)
Generative Ghosts: Anticipating Benefits and Risks of AI Afterlives
von: Morris, Meredith Ringel, et al.
Veröffentlicht: (2024)
von: Morris, Meredith Ringel, et al.
Veröffentlicht: (2024)
Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions
von: Fan, Xianzhe, et al.
Veröffentlicht: (2024)
von: Fan, Xianzhe, et al.
Veröffentlicht: (2024)
Recognising, Anticipating, and Mitigating LLM Pollution of Online Behavioural Research
von: Rilla, Raluca, et al.
Veröffentlicht: (2025)
von: Rilla, Raluca, et al.
Veröffentlicht: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
von: Cheng, Myra, et al.
Veröffentlicht: (2026)
Echoes of Power: Investigating Geopolitical Bias in US and China Large Language Models
von: Pacheco, Andre G. C., et al.
Veröffentlicht: (2025)
von: Pacheco, Andre G. C., et al.
Veröffentlicht: (2025)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
von: Choi, Sooyung, et al.
Veröffentlicht: (2025)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
von: Shieh, Evan, et al.
Veröffentlicht: (2024)
von: Shieh, Evan, et al.
Veröffentlicht: (2024)
AI Consciousness and Public Perceptions: Four Futures
von: Fernandez, Ines, et al.
Veröffentlicht: (2024)
von: Fernandez, Ines, et al.
Veröffentlicht: (2024)
Equity Bias: An Ethical Framework for AI Design
von: Lockwood, Mary
Veröffentlicht: (2026)
von: Lockwood, Mary
Veröffentlicht: (2026)
RealHarm: A Collection of Real-World Language Model Application Failures
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
von: Jeune, Pierre Le, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rolling in the deep of cognitive and AI biases
von: Tantalaki, Nicoleta, et al.
Veröffentlicht: (2024) -
FAIRTOPIA: Envisioning Multi-Agent Guardianship for Disrupting Unfair AI Pipelines
von: Vakali, Athena, et al.
Veröffentlicht: (2025) -
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
von: Mun, Jimin, et al.
Veröffentlicht: (2024) -
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026) -
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
von: Vaccaro, Michelle, et al.
Veröffentlicht: (2026)