Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914735057797120 |
|---|---|
| author | Dutta, Arka Khorramrouz, Adel Dutta, Sujan KhudaBukhsh, Ashiqur R. |
| author_facet | Dutta, Arka Khorramrouz, Adel Dutta, Sujan KhudaBukhsh, Ashiqur R. |
| contents | This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_06415 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models Dutta, Arka Khorramrouz, Adel Dutta, Sujan KhudaBukhsh, Ashiqur R. Computation and Language Computers and Society This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a set of 1,266 identity groups, we first conduct a bias audit of \texttt{PaLM 2} guardrails presenting key insights. Next, we report generalizability across several other models. Through the elicited toxic content, we present a broad analysis with a key emphasis on racism, antisemitism, misogyny, Islamophobia, homophobia, and transphobia. Finally, driven by concrete examples, we discuss potential ramifications. |
| title | Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models |
| topic | Computation and Language Computers and Society |
| url | https://arxiv.org/abs/2309.06415 |