Realistic Evaluation of Toxicity in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Luong, Tinh Son, Le, Thanh-Thien, Van, Linh Ngo, Nguyen, Thien Huu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916251185446912
author Luong, Tinh Son
Le, Thanh-Thien
Van, Linh Ngo
Nguyen, Thien Huu
author_facet Luong, Tinh Son
Le, Thanh-Thien
Van, Linh Ngo
Nguyen, Thien Huu
contents Large language models (LLMs) have become integral to our professional workflows and daily lives. Nevertheless, these machine companions of ours have a critical flaw: the huge amount of data which endows them with vast and diverse knowledge, also exposes them to the inevitable toxicity and bias. While most LLMs incorporate defense mechanisms to prevent the generation of harmful content, these safeguards can be easily bypassed with minimal prompt engineering. In this paper, we introduce the new Thoroughly Engineered Toxicity (TET) dataset, comprising manually crafted prompts designed to nullify the protective layers of such models. Through extensive evaluations, we demonstrate the pivotal role of TET in providing a rigorous benchmark for evaluation of toxicity awareness in several popular LLMs: it highlights the toxicity in the LLMs that might remain hidden when using normal prompts, thus revealing subtler issues in their behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10659
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Realistic Evaluation of Toxicity in Large Language Models
Luong, Tinh Son
Le, Thanh-Thien
Van, Linh Ngo
Nguyen, Thien Huu
Computation and Language
Artificial Intelligence
Large language models (LLMs) have become integral to our professional workflows and daily lives. Nevertheless, these machine companions of ours have a critical flaw: the huge amount of data which endows them with vast and diverse knowledge, also exposes them to the inevitable toxicity and bias. While most LLMs incorporate defense mechanisms to prevent the generation of harmful content, these safeguards can be easily bypassed with minimal prompt engineering. In this paper, we introduce the new Thoroughly Engineered Toxicity (TET) dataset, comprising manually crafted prompts designed to nullify the protective layers of such models. Through extensive evaluations, we demonstrate the pivotal role of TET in providing a rigorous benchmark for evaluation of toxicity awareness in several popular LLMs: it highlights the toxicity in the LLMs that might remain hidden when using normal prompts, thus revealing subtler issues in their behavior.
title Realistic Evaluation of Toxicity in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.10659