Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dutta, Arka, Khorramrouz, Adel, Dutta, Sujan, KhudaBukhsh, Ashiqur R.
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914735057797120
author Dutta, Arka
Khorramrouz, Adel
Dutta, Sujan
KhudaBukhsh, Ashiqur R.
author_facet Dutta, Arka
Khorramrouz, Adel
Dutta, Sujan
KhudaBukhsh, Ashiqur R.
contents This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications.
format Preprint
id arxiv_https___arxiv_org_abs_2309_06415
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
Dutta, Arka
Khorramrouz, Adel
Dutta, Sujan
KhudaBukhsh, Ashiqur R.
Computation and Language
Computers and Society
This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications.
title Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2309.06415