Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kou, Zhiqiang, Chen, Junyang, Cai, Xin-Qiang, Xie, Ming-Kun, Liu, Biao, Wang, Changwei, Feng, Lei, Jia, Yuheng, Niu, Gang, Sugiyama, Masashi, Geng, Xin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914097903173632
author Kou, Zhiqiang
Chen, Junyang
Cai, Xin-Qiang
Xie, Ming-Kun
Liu, Biao
Wang, Changwei
Feng, Lei
Jia, Yuheng
Niu, Gang
Sugiyama, Masashi
Geng, Xin
author_facet Kou, Zhiqiang
Chen, Junyang
Cai, Xin-Qiang
Xie, Ming-Kun
Liu, Biao
Wang, Changwei
Feng, Lei
Jia, Yuheng
Niu, Gang
Sugiyama, Masashi
Geng, Xin
contents Large language models (LLMs) have achieved impressive results across a range of natural language processing tasks, but their potential to generate harmful content has raised serious safety concerns. Current toxicity detectors primarily rely on single-label benchmarks, which cannot adequately capture the inherently ambiguous and multi-dimensional nature of real-world toxic prompts. This limitation results in biased evaluations, including missed toxic detections and false positives, undermining the reliability of existing detectors. Additionally, gathering comprehensive multi-label annotations across fine-grained toxicity categories is prohibitively costly, further hindering effective evaluation and development. To tackle these issues, we introduce three novel multi-label benchmarks for toxicity detection: \textbf{Q-A-MLL}, \textbf{R-A-MLL}, and \textbf{H-X-MLL}, derived from public toxicity datasets and annotated according to a detailed 15-category taxonomy. We further provide a theoretical proof that, on our released datasets, training with pseudo-labels yields better performance than directly learning from single-label supervision. In addition, we develop a pseudo-label-based toxicity detection method. Extensive experimental results show that our approach significantly surpasses advanced baselines, including GPT-4o and DeepSeek, thus enabling more accurate and reliable evaluation of multi-label toxicity in LLM-generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15007
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
Kou, Zhiqiang
Chen, Junyang
Cai, Xin-Qiang
Xie, Ming-Kun
Liu, Biao
Wang, Changwei
Feng, Lei
Jia, Yuheng
Niu, Gang
Sugiyama, Masashi
Geng, Xin
Computation and Language
Artificial Intelligence
Large language models (LLMs) have achieved impressive results across a range of natural language processing tasks, but their potential to generate harmful content has raised serious safety concerns. Current toxicity detectors primarily rely on single-label benchmarks, which cannot adequately capture the inherently ambiguous and multi-dimensional nature of real-world toxic prompts. This limitation results in biased evaluations, including missed toxic detections and false positives, undermining the reliability of existing detectors. Additionally, gathering comprehensive multi-label annotations across fine-grained toxicity categories is prohibitively costly, further hindering effective evaluation and development. To tackle these issues, we introduce three novel multi-label benchmarks for toxicity detection: \textbf{Q-A-MLL}, \textbf{R-A-MLL}, and \textbf{H-X-MLL}, derived from public toxicity datasets and annotated according to a detailed 15-category taxonomy. We further provide a theoretical proof that, on our released datasets, training with pseudo-labels yields better performance than directly learning from single-label supervision. In addition, we develop a pseudo-label-based toxicity detection method. Extensive experimental results show that our approach significantly surpasses advanced baselines, including GPT-4o and DeepSeek, thus enabling more accurate and reliable evaluation of multi-label toxicity in LLM-generated content.
title Rethinking Toxicity Evaluation in Large Language Models: A Multi-Label Perspective
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.15007