ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Boyang, Shou, Hongzhe, Liang, Yuanyuan, Zhang, Jingbin, Zhou, Fang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917406278942720
author Li, Boyang
Shou, Hongzhe
Liang, Yuanyuan
Zhang, Jingbin
Zhou, Fang
author_facet Li, Boyang
Shou, Hongzhe
Liang, Yuanyuan
Zhang, Jingbin
Zhou, Fang
contents Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evidence spans. We propose \textbf{ToxiTrace}, an explainability-oriented method for BERT-style encoders with three components: (1) \textbf{CuSA}, which refines encoder-derived saliency cues into fine-grained toxic spans with lightweight LLM guidance; (2) \textbf{GCLoss}, a gradient-constrained objective that concentrates token-level saliency on toxic evidence while suppressing irrelevant activations; and (3) \textbf{ARCL}, which constructs sample-specific contrastive reasoning pairs to sharpen the semantic boundary between toxic and non-toxic content. Experiments show that ToxiTrace improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent, human-readable explanations. We have released the model at https://huggingface.co/ArdLi/ToxiTrace.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection
Li, Boyang
Shou, Hongzhe
Liang, Yuanyuan
Zhang, Jingbin
Zhou, Fang
Computation and Language
Existing Chinese toxic content detection methods mainly target sentence-level classification but often fail to provide readable and contiguous toxic evidence spans. We propose \textbf{ToxiTrace}, an explainability-oriented method for BERT-style encoders with three components: (1) \textbf{CuSA}, which refines encoder-derived saliency cues into fine-grained toxic spans with lightweight LLM guidance; (2) \textbf{GCLoss}, a gradient-constrained objective that concentrates token-level saliency on toxic evidence while suppressing irrelevant activations; and (3) \textbf{ARCL}, which constructs sample-specific contrastive reasoning pairs to sharpen the semantic boundary between toxic and non-toxic content. Experiments show that ToxiTrace improves classification accuracy and toxic span extraction while preserving efficient encoder-based inference and producing more coherent, human-readable explanations. We have released the model at https://huggingface.co/ArdLi/ToxiTrace.
title ToxiTrace: Gradient-Aligned Training for Explainable Chinese Toxicity Detection
topic Computation and Language
url https://arxiv.org/abs/2604.12321