RealHarm: A Collection of Real-World Language Model Application Failures

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Jeune, Pierre Le, Liu, Jiaen, Rossi, Luca, Dora, Matteo
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915484960555008
author Jeune, Pierre Le
Liu, Jiaen
Rossi, Luca
Dora, Matteo
author_facet Jeune, Pierre Le
Liu, Jiaen
Rossi, Luca
Dora, Matteo
contents Language model deployments in consumer-facing applications introduce numerous risks. While existing research on harms and hazards of such applications follows top-down approaches derived from regulatory frameworks and theoretical analyses, empirical evidence of real-world failure modes remains underexplored. In this work, we introduce RealHarm, a dataset of annotated problematic interactions with AI agents built from a systematic review of publicly reported incidents. Analyzing harms, causes, and hazards specifically from the deployer's perspective, we find that reputational damage constitutes the predominant organizational harm, while misinformation emerges as the most common hazard category. We empirically evaluate state-of-the-art guardrails and content moderation systems to probe whether such systems would have prevented the incidents, revealing a significant gap in the protection of AI applications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealHarm: A Collection of Real-World Language Model Application Failures
Jeune, Pierre Le
Liu, Jiaen
Rossi, Luca
Dora, Matteo
Computers and Society
Artificial Intelligence
Computation and Language
Cryptography and Security
Language model deployments in consumer-facing applications introduce numerous risks. While existing research on harms and hazards of such applications follows top-down approaches derived from regulatory frameworks and theoretical analyses, empirical evidence of real-world failure modes remains underexplored. In this work, we introduce RealHarm, a dataset of annotated problematic interactions with AI agents built from a systematic review of publicly reported incidents. Analyzing harms, causes, and hazards specifically from the deployer's perspective, we find that reputational damage constitutes the predominant organizational harm, while misinformation emerges as the most common hazard category. We empirically evaluate state-of-the-art guardrails and content moderation systems to probe whether such systems would have prevented the incidents, revealing a significant gap in the protection of AI applications.
title RealHarm: A Collection of Real-World Language Model Application Failures
topic Computers and Society
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2504.10277