ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Kangwei, Cheng, Siyuan, Tian, Bozhong, Liang, Xiaozhuan, Yin, Yuyang, Han, Meng, Zhang, Ningyu, Hooi, Bryan, Chen, Xi, Deng, Shumin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918124202229760
author Liu, Kangwei
Cheng, Siyuan
Tian, Bozhong
Liang, Xiaozhuan
Yin, Yuyang
Han, Meng
Zhang, Ningyu
Hooi, Bryan
Chen, Xi
Deng, Shumin
author_facet Liu, Kangwei
Cheng, Siyuan
Tian, Bozhong
Liang, Xiaozhuan
Yin, Yuyang
Han, Meng
Zhang, Ningyu
Hooi, Bryan
Chen, Xi
Deng, Shumin
contents Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However, existing resources for harmful content detection are predominantly focused on English, with Chinese datasets remaining scarce and often limited in scope. We present a comprehensive, professionally annotated benchmark for Chinese content harm detection, which covers six representative categories and is constructed entirely from real-world data. Our annotation process further yields a knowledge rule base that provides explicit expert knowledge to assist LLMs in Chinese harmful content detection. In addition, we propose a knowledge-augmented baseline that integrates both human-annotated knowledge rules and implicit knowledge from large language models, enabling smaller models to achieve performance comparable to state-of-the-art LLMs. Code and data are available at https://github.com/zjunlp/ChineseHarm-bench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10960
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
Liu, Kangwei
Cheng, Siyuan
Tian, Bozhong
Liang, Xiaozhuan
Yin, Yuyang
Han, Meng
Zhang, Ningyu
Hooi, Bryan
Chen, Xi
Deng, Shumin
Computation and Language
Artificial Intelligence
Cryptography and Security
Information Retrieval
Machine Learning
Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However, existing resources for harmful content detection are predominantly focused on English, with Chinese datasets remaining scarce and often limited in scope. We present a comprehensive, professionally annotated benchmark for Chinese content harm detection, which covers six representative categories and is constructed entirely from real-world data. Our annotation process further yields a knowledge rule base that provides explicit expert knowledge to assist LLMs in Chinese harmful content detection. In addition, we propose a knowledge-augmented baseline that integrates both human-annotated knowledge rules and implicit knowledge from large language models, enabling smaller models to achieve performance comparable to state-of-the-art LLMs. Code and data are available at https://github.com/zjunlp/ChineseHarm-bench.
title ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2506.10960