Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xintong, Liu, Yixiao, Pan, Jingheng, Ding, Liang, Wang, Longyue, Biemann, Chris
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912385576468480
author Wang, Xintong
Liu, Yixiao
Pan, Jingheng
Ding, Liang
Wang, Longyue
Biemann, Chris
author_facet Wang, Xintong
Liu, Yixiao
Pan, Jingheng
Ding, Liang
Wang, Longyue
Biemann, Chris
contents Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent. This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context. We present ToxiRewriteCN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity. The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans. It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues. We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity. Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs. We release ToxiRewriteCN to support future research on controllable, sentiment-aware detoxification for Chinese.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
Wang, Xintong
Liu, Yixiao
Pan, Jingheng
Ding, Liang
Wang, Longyue
Biemann, Chris
Computation and Language
Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent. This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context. We present ToxiRewriteCN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity. The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans. It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues. We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity. Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs. We release ToxiRewriteCN to support future research on controllable, sentiment-aware detoxification for Chinese.
title Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
topic Computation and Language
url https://arxiv.org/abs/2505.15297