Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Yibo, Zhu, Jiapeng, Xu, Can, Liu, Yao, Li, Xiang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915316429225984
author Zhao, Yibo
Zhu, Jiapeng
Xu, Can
Liu, Yao
Li, Xiang
author_facet Zhao, Yibo
Zhu, Jiapeng
Xu, Can
Liu, Yao
Li, Xiang
contents The rapid growth of social media platforms has raised significant concerns regarding online content toxicity. When Large Language Models (LLMs) are used for toxicity detection, two key challenges emerge: 1) the absence of domain-specific toxic knowledge leads to false negatives; 2) the excessive sensitivity of LLMs to toxic speech results in false positives, limiting freedom of speech. To address these issues, we propose a novel method called MetaTox, leveraging graph search on a meta-toxic knowledge graph to enhance hatred and toxicity detection. First, we construct a comprehensive meta-toxic knowledge graph by utilizing LLMs to extract toxic information through a three-step pipeline, with toxic benchmark datasets serving as corpora. Second, we query the graph via retrieval and ranking processes to supplement accurate, relevant toxic knowledge. Extensive experiments and in-depth case studies across multiple datasets demonstrate that our MetaTox significantly decreases the false positive rate while boosting overall toxicity detection performance. Our code is available at https://github.com/YiboZhao624/MetaTox.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15268
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
Zhao, Yibo
Zhu, Jiapeng
Xu, Can
Liu, Yao
Li, Xiang
Computation and Language
Artificial Intelligence
The rapid growth of social media platforms has raised significant concerns regarding online content toxicity. When Large Language Models (LLMs) are used for toxicity detection, two key challenges emerge: 1) the absence of domain-specific toxic knowledge leads to false negatives; 2) the excessive sensitivity of LLMs to toxic speech results in false positives, limiting freedom of speech. To address these issues, we propose a novel method called MetaTox, leveraging graph search on a meta-toxic knowledge graph to enhance hatred and toxicity detection. First, we construct a comprehensive meta-toxic knowledge graph by utilizing LLMs to extract toxic information through a three-step pipeline, with toxic benchmark datasets serving as corpora. Second, we query the graph via retrieval and ranking processes to supplement accurate, relevant toxic knowledge. Extensive experiments and in-depth case studies across multiple datasets demonstrate that our MetaTox significantly decreases the false positive rate while boosting overall toxicity detection performance. Our code is available at https://github.com/YiboZhao624/MetaTox.
title Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.15268