Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Magu, Rijul, Dutta, Arka, Kim, Sean, KhudaBukhsh, Ashiqur R., De Choudhury, Munmun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910003868205056
author Magu, Rijul
Dutta, Arka
Kim, Sean
KhudaBukhsh, Ashiqur R.
De Choudhury, Munmun
author_facet Magu, Rijul
Dutta, Arka
Kim, Sean
KhudaBukhsh, Ashiqur R.
De Choudhury, Munmun
contents Large Language Models (LLMs) have been shown to demonstrate imbalanced biases against certain groups. However, the study of unprovoked targeted attacks by LLMs towards at-risk populations remains underexplored. Our paper presents three novel contributions: (1) the explicit evaluation of LLM-generated attacks on highly vulnerable mental health groups; (2) a network-based framework to study the propagation of relative biases; and (3) an assessment of the relative degree of stigmatization that emerges from these attacks. Our analysis of a recently released large-scale bias audit dataset reveals that mental health entities occupy central positions within attack narrative networks, as revealed by a significantly higher mean centrality of closeness (p-value = 4.06e-10) and dense clustering (Gini coefficient = 0.7). Drawing from an established stigmatization framework, our analysis indicates increased labeling components for mental health disorder-related targets relative to initial targets in generation chains. Taken together, these insights shed light on the structural predilections of large language models to heighten harmful discourse and highlight the need for suitable approaches for mitigation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
Magu, Rijul
Dutta, Arka
Kim, Sean
KhudaBukhsh, Ashiqur R.
De Choudhury, Munmun
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
Social and Information Networks
J.4; K.4.1; K.4.2
Large Language Models (LLMs) have been shown to demonstrate imbalanced biases against certain groups. However, the study of unprovoked targeted attacks by LLMs towards at-risk populations remains underexplored. Our paper presents three novel contributions: (1) the explicit evaluation of LLM-generated attacks on highly vulnerable mental health groups; (2) a network-based framework to study the propagation of relative biases; and (3) an assessment of the relative degree of stigmatization that emerges from these attacks. Our analysis of a recently released large-scale bias audit dataset reveals that mental health entities occupy central positions within attack narrative networks, as revealed by a significantly higher mean centrality of closeness (p-value = 4.06e-10) and dense clustering (Gini coefficient = 0.7). Drawing from an established stigmatization framework, our analysis indicates increased labeling components for mental health disorder-related targets relative to initial targets in generation chains. Taken together, these insights shed light on the structural predilections of large language models to heighten harmful discourse and highlight the need for suitable approaches for mitigation.
title Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
Social and Information Networks
J.4; K.4.1; K.4.2
url https://arxiv.org/abs/2504.06160